✳ A LITTLE TIME TRAVEL FOR YOUR WORDS

Back to the beginning.

Before the conversations. Before the billions of parameters.
Meet GPT-1, the model that started it all. Give it a few words.
See where its imagination goes.

Let’s try the original
117MPARAMETERS
12TRANSFORMER LAYERS
512TOKEN CONTEXT
Jun ’18THE FIRST CHAPTER
Original weights.
A fresh blank page.
01 /

The playground

Ready to start
Start GPT-1 on your device.

Checking first download size…

YOUR PROMPT
A sentence, a story, a small spark.Counting tokens…
NEED A FIRST LINE?
⌘ ↵
Experimental controlsA few ways to bend the next word.
Off

Keep only the k most likely tokens. Zero lets top-p choose.

1.10

Discourage tokens already in memory. 1.00 turns it off.

Off

Prevent a token sequence of this size from appearing again.

Reuse a seed and settings to repeat an experiment.

Pauses 0.5 seconds on each token so you can read its candidates and probabilities. Some words have several tokens.

✳

These controls change sampling.
The model is still the original GPT-1.

The continuation

WAITING FOR A SPARK

A blank page. An early mind.

Your words go in. A little 2018 imagination comes out.

↳ Think of it as a writing partner from 2018. It continues your text and is best at English fiction.

02 / A NOTE FROM THE ARCHIVE

Small model.
Big first step.

In 2018, GPT-1 explored a simple idea: learn from books, then use that understanding for other language tasks. Twelve transformer layers. About 117 million parameters. The beginning of a very long story.

Expect wandering sentences, lowercase text, and the occasional plot twist. This early model wasn’t trained to follow instructions, and its output can be inaccurate or biased.

Read the original paper
FROM THE ARCHIVE / 2018

The first GPT.

GPT stands for Generative Pre-trained Transformer. This playground runs the original openai-community/openai-gpt checkpoint from Hugging Face directly in your browser.

It predicts the next token in your text. It has no chat training, so starting a story usually works better than asking a question.

Architecture
12 layers · 12 attention heads
Training data
BooksCorpus
Context window
512 tokens, including your prompt
Long-form output
Up to 2,048 tokens per run
Privacy
Your prompts stay on this device.

Long-form mode keeps the most recent 512 tokens in memory and forgets older text. It can write longer passages but may lose the original thread. The original checkpoint is compressed to four-bit weights for browser use. This can slightly change its output. GPU and CPU results may differ; seeds repeat within the same backend. Generated text may be inaccurate, biased, or offensive.

Explore the model card