Academy · Build an LLM From Scratch

Step 1 of 1 · 50 min

A line is a model

You already learned the core equation of machine learning in school. Nobody told you that's what it was — and it's the same equation, at any scale, all the way up to GPT.

A line is a model

You already learned the core equation of machine learning in school. Nobody told you that's what it was. Let's fix that.

Ground zero: what a graph even is

Skip this section if slope, intercept and coordinates already mean something to you — nothing later needs more than what's restated as it comes up.

Suppose I tell you: a student studied 4 hours and scored 52 marks. That's two numbers about one student. Fifty students would be a hundred numbers — an unreadable wall of digits.

A graph turns pairs of numbers into a picture your eyes can read instantly. Draw two lines that meet in a corner: the flat one going across is the x-axis, the upright one is the y-axis, and where they meet — where both numbers are zero — is the origin. Any pair of numbers becomes a place on the picture: the first number says how far to walk right, the second says how far to climb up. Put a dot where you land. A point is written (x, y) — always across first, then up — so (4, 52) means "4 across, 52 up". That pair is called its coordinates.

Plot fifty students as fifty dots, and if studying more genuinely produces higher marks, the dots drift upward as you move right — visible in half a second, no arithmetic required. If the dots roughly follow a straight path, you can draw a straight line through them. That line is a summary of all fifty students, and — more usefully — a prediction machine for a student you haven't met: pick an hours-studied value on the bottom, go up to the line, read the marks across. That's a forecast.

So the only question left is: how do you describe one particular straight line, out of the infinitely many that exist? It turns out you only need two facts.

Fact 1 — the tilt ("slope"). How steeply does the line climb? Take one step to the right and see how far the line went up. Step 1 right and the line rises by 3 → the tilt is 3. Rises by 0.5 → a gentle line. Goes down by 2 → the tilt is −2; negative just means downhill. That number is the slope, and the only definition worth keeping is a sentence, not a formula: slope answers "if x goes up by 1, how much does y change?"

Fact 2 — the starting height ("intercept"). Two lines can share a tilt but sit at different heights, like two parallel roads on a hillside — so tilt alone isn't enough. The convention is to record the line's height when x = 0, where it crosses the upright y-axis. That number is the intercept — literally "the crossing point".

The interactive below lets you feel this: move only the tilt slider and watch the line pivot around a fixed crossing point; then move only the height slider and watch it slide up and down with the tilt untouched. Two independent controls.

Put the two facts together and you have everything: to find y for any x, start at the crossing height, then climb by the tilt, once for every step of x. Tilt 3, crossing height 10, and x = 4: climb 3 four times (12) from a start of 10, so y = 22. As a recipe:

y = (tilt × x) + height          — the plain-English version

Mathematicians got tired of writing "tilt" and "height", so tilt became m and height became c:

y = mx + c                        — the same recipe, abbreviated

That's the whole equation. It isn't a rule handed down from above — it's "start at the height, climb by the tilt" in shorthand. Read it back as that sentence every time you see it.

The thing you already know

So here's where we've landed: y = mx + c, where m is the slope (the tilt) and c is the intercept (the crossing height).

Whether you just built that above or met it years ago in school, it probably felt like a rule for passing an exam. Here's the part nobody says out loud: that equation is a machine that makes predictions. Feed it an x, it hands back a y. That's all any AI model does — including the one you're reading this from. GPT takes in text and hands back the next word. Different scale, identical shape.

The only difference between y = mx + c and a large language model is how many m's there are — plus one extra ingredient met in the next lesson but one. Your school line has one m. GPT-4 has roughly a trillion. Set the extra ingredient aside for now; the mechanics genuinely don't change.

Machine learning people renamed the two numbers: m becomes the weight, c becomes the bias:

y = wx + b

Same equation, new clothes — this is the notation you'll see in every paper and every line of PyTorch from here on.

Weight is a good name: it answers how much does this input matter? A big w means x has heavy influence. A w near zero means "ignore this input, it's noise." A negative w means "the more of this, the worse the outcome." Bias is a worse name and confuses everyone — it has nothing to do with prejudice. It's just the model's starting assumption before it looks at any input: if x is zero, what's your answer? That's b. Think of it as the model's default mood.

Your first model, by hand

Eight students. For each one we know hours studied and the marks they got. We want a model that predicts marks from hours, so that for a new student who studied 5.5 hours we can guess their score.

Hours studiedMarks
122
231
338
452
557
668
772
885

The interactive below lets you drag w (the weight/slope) and b (the bias/intercept) until the line runs through the cloud of dots as well as you can manage — trust your eyes, the meaning of the numbers is unpacked just below. Try to get the loss under 30 using only the sliders, then hit "let the machine do it" and watch it find a better line than you did, the same way, just faster. Notice how it feels while you do it by hand: nudge w, loss drops; nudge it more, loss climbs again. That feeling is gradient descent.

What "error" actually means

You were adjusting sliders to make the line "look right". A computer has no eyes — right has to be a number, or there's nothing to improve. Three steps:

Step 1 — Predict. For the student who studied 1 hour, the model says ŷ = w(1) + b. The little hat on ŷ means "predicted", to distinguish it from the real y.

Step 2 — Measure the miss. The real mark was 22; the model said something else. The gap is y − ŷ — one miss per student.

Step 3 — Add up all the misses, with a trap to avoid: if you just add the gaps, a model that overshoots one student by +30 and undershoots another by −30 scores a total error of zero. Terrible model, perfect score — the positives and negatives cancel and hide the damage.

So each gap is squared before adding, which does two deliberate things: it kills the sign (−30 and +30 both become 900, so nothing cancels), and it punishes big misses harder (off by 10 costs 100; off by 20 costs 400 — four times worse, not twice, which pushes the model to fix its worst mistakes first). Divide by the number of students so the number means "average miss per student" rather than "bigger dataset, bigger number", and you get:

L = (1/n) · Σ (y − ŷ)2      — Mean Squared Error

Read left to right in plain English: take every student, find how far off you were, square it, average them. That Σ just means "add up all of them". This number is the loss — lower is better, zero is perfect, and training a model, any model including GPT, is nothing but making this one number go down. Not "learn patterns", not "understand language" — find the knob settings that minimise one number. Everything else is engineering around that sentence.

The landscape

The second panel in the interactive above plots every possible pair of settings — w across, b up — coloured by the loss you'd get with those settings: dark means low loss (good), bright means high loss (bad). It's a valley with exactly one lowest point — the best line that exists for this data. Your dot sits wherever your sliders currently are.

Nobody has to know where the bottom is. If you're standing on a slope in fog, you can still find your way down: feel which direction tilts downhill, take a step, repeat, and you'll reach the bottom without ever seeing it. That's gradient descent, and it's how every model on Earth is trained. "Feel the slope" is the derivative, which is the next lesson — and it's much less scary than school made it look.

This was a neuron the whole time

Here's the payoff. What you just built, drawn the way a textbook draws it, sits inside the interactive above: one input, multiplied by a weight, with a bias added, producing an output. That diagram isn't an analogy — it's the same equation, redrawn. A neuron is wx + b, with one small thing wrapped around it that the non-linearity lesson adds. When people say a network has "a billion parameters", they mean a billion of these w's and b's — not a billion clever ideas. A billion numbers, each one a slider, all being nudged to make one loss number go down.

Real neurons take many inputs. Suppose marks depend on hours studied and hours slept — every input gets its own weight, because every input matters a different amount:

ŷ = w1x1 + w2x2 + b        — two inputs, two weights, still one bias

And that's the pattern all the way up: ten inputs, ten weights; fifty thousand inputs, fifty thousand weights. The equation never changes shape, it just gets longer. Later this gets written as W·x + b where W is a matrix — shorthand so nobody has to type a million plus signs. Nothing new is happening.

The same thing, in code

Everything in this lesson, in plain Python, no libraries. Read it and check that each line matches something you now understand:

# our eight students: (hours studied, marks)
data = [(1,22), (2,31), (3,38), (4,52),
        (5,57), (6,68), (7,72), (8,85)]

def predict(x, w, b):
    """The model. This one line is the whole thing."""
    return w * x + b

def loss(w, b):
    """How wrong are we, on average, squared?"""
    total = 0
    for x, y in data:
        gap = y - predict(x, w, b)     # the miss
        total += gap ** 2              # square it: no cancelling, big misses hurt more
    return total / len(data)           # average per student

# try a bad guess vs a good one
print(loss(2.0, 40.0))    # → 264.875  ugly (floating-point noise on the tail)
print(loss(8.82, 13.43))  # → 3.567...  the best line that exists

Nine lines contain a complete machine learning model, a prediction function, and a loss function. What's missing is the part that finds 8.82 on its own instead of you typing it — that's the next lesson.

Check yourself

Answer before revealing the explanation. Getting one wrong is useful information.

Tilt and height — drag one, then the other
Fit the line by hand — and watch it as a neuron

Practice

1. A line has slope 4. What does that tell you?

2. A line is written y = 3x + 7. What is y when x = 0?

3. A model predicting house price from size has weight w = 0. What does that tell you?

4. Why square the errors instead of just taking their absolute value?

5. On the landscape panel, what does one single point represent?

6. What is genuinely different between this line and GPT?