The “Conversation” Is Just Tokens on Repeat
Large Language Models (LLMs) generate text by predicting one token at a time, repeatedly, until they complete a response. This mechanism differs fundamentally from human comprehension, yet enables surprisingly sophisticated outputs. The “conversation” isn't true understanding; it’s a probabilistic dance of statistical associations.
Before processing, input text undergoes tokenization, where it's broken into common words and subword chunks. For example, "unbelievable" might split into "un", "believe", and "able." A token averages about three-quarters of a word. Token counts directly impact API costs, context window limits, and sometimes latency. Explore this process with OpenAI’s Tokenizer or Hugging Face’s tokenizer summary.
LLMs are inherently stateless; they possess no intrinsic memory between requests. When a chat interface appears to "remember" past interactions, it achieves this by resending the entire conversation history (or relevant stored facts) with each new query. This consumes valuable context space and directly contributes to token usage.
Inside the Transformer: Weights, Attention, and Training
Beneath the surface of token prediction lies a complex architecture. Think of it like calculating a house price: input the size (1,000 sq ft), multiply by a weight (say, 300), and get an output ($300,000). LLMs scale this up, stacking hundreds of layers with thousands of "neurons" to process far more intricate patterns. These billions of weights, collectively called parameters, are fixed during inference but adjusted during training.
The "T" in GPT stands for Transformer, and its key innovation is attention. Imagine the word "bank": its meaning changes dramatically between "river bank" and "financial bank." Attention allows each token to dynamically weigh the importance of other tokens in the input sequence, disambiguating meaning by linking "bank" to "river" or "money." This mechanism, introduced in the seminal "Attention Is All You Need" paper, prevents tokens from being processed in isolation.
The learning pipeline involves several stages. Pre-training begins with random weights, where the model predicts missing tokens and adjusts its weights via backpropagation based on error. This iterative process, repeated trillions of times, refines predictions. Fine-tuning then adapts the pre-trained model to specific tasks using smaller, curated datasets. Finally, reinforcement learning can further optimize behavior by scoring model outputs and adjusting weights to favor high-scoring responses, often incorporating human feedback (RLHF) or automated checks.
Make Models Smaller—and Give Them Fresh Facts
Adapting large models for specific tasks can be resource-intensive, but techniques like LoRA (Low-Rank Adaptation) offer a cost-effective solution. LoRA freezes the vast majority of a model’s original weights, training only a small, additional set of parameters on top—often less than 1% of the base model's size. This significantly reduces the computational overhead, enabling fine-tuning on a single GPU.
Quantization addresses the memory footprint of these colossal models. Each weight in an LLM typically stores as a 16-bit number. An 8-billion-parameter model, for example, consumes roughly 16 gigabytes. Quantization reduces this by storing weights with fewer bits (e.g., 8 or 4 bits), rounding them to less precise values. While this shrinks the model—a 4-bit quantized version of the same model might be around 5 gigabytes, runnable on a laptop—it introduces a trade-off, potentially sacrificing some accuracy.
To extend a model's knowledge beyond its training data, Retrieval-Augmented Generation (RAG) leverages external information. This process begins by converting text into embeddings—long lists of numbers that numerically capture semantic meaning, enabling comparison by similarity. These embeddings, along with their original text, are then stored in a vector database.
When a user poses a question, the application first transforms that query into an embedding. It then searches the vector database for semantically similar passages. These retrieved passages are then appended to the original prompt, allowing the LLM to generate an answer informed by current or proprietary data it was never explicitly trained on. Explore how text converts to tokens and embeddings with the OpenAI Tokenizer Tool.
Enjoying this? Get one like it in your inbox each morning.
one email a day · unsubscribe in two clicks · no third-party tracking
From One Reply to Agents That Take Action
Models, in their essence, are text generators. They cannot independently search the web or execute commands. This is where tool calling comes in: the model identifies a need for an external function, then requests a specific, allowed action in a structured format. Application code, not the model itself, receives this request, executes the tool, and returns the result to the model as plain text.
Many services expose these tools through the Model-Client Protocol (MCP), a common connection pattern that allows compatible AI applications to discover and interact with external capabilities. MCP defines how an application can communicate with a tool, but it does not dictate the model's internal reasoning or capabilities. The model uses MCP to request actions; it does not run them.
These concepts converge in an agent loop. An AI agent receives a task, identifies necessary tools, and calls them via MCP. It then inspects the tool's output, updates its internal state, and repeats the process until the task is complete. This iterative cycle enables complex workflows, from booking flights to analyzing data.
While powerful, these autonomous workflows require careful management. Always validate tool outputs, configure appropriate permissions to prevent unauthorized access, and carefully consider data exposure. Agents excel at structured tasks, but their autonomy has limits; human oversight remains crucial for robust, safe operation.
Frequently Asked Questions
What is an LLM in simple terms?
A large language model predicts likely next tokens from its input, repeating the process to generate a response.
Do AI chatbots remember past conversations?
The model itself is generally stateless. A chat app can create continuity by sending conversation history or saved information with a new request.
What is the difference between RAG and fine-tuning?
RAG retrieves relevant information at response time and adds it to the prompt. Fine-tuning adjusts model weights using additional training examples.
How is an AI agent different from a chatbot?
An agent can run a loop of reasoning and actions, using tools, checking results, and continuing toward a task rather than only replying once.

