Back to the beginning.
Before the conversations. Before the billions of parameters.
Meet GPT-1, the model that started it all. Give it a few words.
See where its imagination goes.
A fresh blank page.
The playground
Checking first download size…
Experimental controlsA few ways to bend the next word.
Keep only the k most likely tokens. Zero lets top-p choose.
Discourage tokens already in memory. 1.00 turns it off.
Prevent a token sequence of this size from appearing again.
Reuse a seed and settings to repeat an experiment.
Pauses 0.5 seconds on each token so you can read its candidates and probabilities. Some words have several tokens.
These controls change sampling.
The model is still the original GPT-1.
The continuation
WAITING FOR A SPARKA blank page. An early mind.
Your words go in. A little 2018 imagination comes out.
The next-word lottery.
Generate a passage to see the latest next-token candidates.
Memory: 0 / 512 tokens
Waiting for the first token…
↳ Think of it as a writing partner from 2018. It continues your text and is best at English fiction.
Small model.
Big first step.
In 2018, GPT-1 explored a simple idea: learn from books, then use that understanding for other language tasks. Twelve transformer layers. About 117 million parameters. The beginning of a very long story.
Expect wandering sentences, lowercase text, and the occasional plot twist. This early model wasn’t trained to follow instructions, and its output can be inaccurate or biased.
Read the original paper