How can next-token predictors talk?
From phone keyboard suggestions to conversations about the meaning of life
I. “LLMs Are Autocomplete”
There is virtually infinite content on the web reminding us that LLMs are “glorified autocomplete”, or “just next token predictors”. What they don’t explain and has been consuming me for quite some time is how this leads to something that can have a conversation. It’s deeply counterintuitive. Autocompletes are good for things like search bars, phone keyboard suggestions or filling your credit card info. Now they talk?
When I started using OpenAI’s API, it got even more confusing. Now I was sending a full-fledged JSON that had keys — system, user, and assistant — with different meanings associated to them. These things can handle JSON with this level of abstraction? It was time to unpack.
# Sample messages array that you could send to any LLM
messages = [
{"role": "system", "content": "You are a helpful assistant"},
{"role": "user", "content": "Explain me who you are"},
{"role": "assistant", "content": "I am the LLM"},
{"role": "user", "content": "What does that mean?"},
]
II. A Quick Peek At The Inputs of Lllama, Phi and DeepSeek
The best way to understand something is to take it apart.
My goal was to look at the LLMs inputs and figure out what kind of data they were receiving in first place. Is it really JSON? In order to do that, I’ve setup a HuggingFace account and selected some famous OpenSource models to analyze: DeepSeek-V3.1, Llama-3.1-8B by Meta and Phi-4-mini-instruct by Microsoft. Just before sending the prompt, I’ve printed their content to see exactly what they are. The contents are below, but I recommend you get the Jupyter Notebook), test, and see it for yourself.
=== microsoft/Phi-4-mini-instruct ===
<|system|>You are a helpful assistant<|end|>
<|user|>Explain me who you are<|end|>
<|assistant|>I am the LLM<|end|>
<|user|>What does that mean?<|end|>
<|assistant|>
=== deepseek-ai/DeepSeek-V3.1 ===
<|begin▁of▁sentence|>You are a helpful assistant
<|User|>Explain me who you are
<|Assistant|></think>I am the LLM<|end▁of▁sentence|>
<|User|>What does that mean?
<|Assistant|></think>
=== meta-llama/Meta-Llama-3.1-8B-Instruct ===
<|begin_of_text|><|start_header_id|>system<|end_header_id|>
Cutting Knowledge Date: December 2023
Today Date: 26 Jul 2024
You are a helpful assistant<|eot_id|><|start_header_id|>user<|end_header_id|>
Explain me who you are<|eot_id|><|start_header_id|>assistant<|end_header_id|>
I am the LLM<|eot_id|><|start_header_id|>user<|end_header_id|>
What does that mean?<|eot_id|><|start_header_id|>assistant<|end_header_id|>
So is this what LLMs really receive? The results were eye-opening. What we see is a document that contains the whole conversation and looks very much like an HMTL file. The tags like <|user|> and <|system|> are called special tokens and much like semantic HTML they help attribute meaning to the content. Phi’s example is easier to understand due to its simplicity. You can see that the document ends with the <|assistant|>, which is exactly where the LLM is supposed to insert the next answer in the conversation, which means, in practice, autocompleting the document.
III. But How Does it Know What System, Assistant and User Are?
One important thing to notice is that LLMs do not understand natively what system, assistant and user are. The base model, which is the model that was trained with massive amounts of text, has no idea that system defines its personality, user carries someone's message, and assistant is its own voice.
There is no coded native functionality that handles this.
What happens is that during fine-tuning the base model is trained with thousands of conversations with the same system→user→assistant structure and, with enough repetition, it learns the pattern. The model does not “understand” that it is an assistant. What it does is identifying the input format and predicting that the most likely next token is an assistant answer. We are at the point now where those predictions are so good for frontier models that we get the illusion of a conversation. But that’s what it is. An illusion.
How can next-token predictors talk? They can’t. We structure text in a way that makes autocomplete produce a conversation. The magic is in the input structure.