Skip to content
AI & LLMs7 min read

What Actually Happens When You Send a Message to ChatGPT?

You type a question, hit enter, and a few seconds later you have an answer that reads like a human wrote it. It feels a little like magic. It isn't. It's math, at a scale that's hard to picture…

  • chatgpt
  • openai
  • gpt
  • Prompt
  • Tokenization
  • ChaiCode
  • engineering
  • AI

You type a question, hit enter, and a few seconds later you have an answer that reads like a human wrote it. It feels a little like magic. It isn't. It's math, at a scale that's hard to picture, running on an architecture that changed AI completely. Let's break down what's actually happening under the hood.

What is an LLM?

LLM stands for Large Language Model, a model trained on a large amount of language data, built by ML engineers who feed it huge volumes of text so it can learn the patterns in how we write and speak.

Here's the problem LLMs actually solve. Before this generation of models, getting a computer to work with language meant writing rigid rules, if this word appears do this, match this pattern do that. That approach breaks the moment language does anything even slightly unexpected, which is most of the time. LLMs don't work off fixed rules. They learn statistical patterns from data, which means they generalize to sentences they've never seen before.

You already know the popular examples: GPT-4o from OpenAI, Claude from Anthropic, Gemini from Google, Llama from Meta.

One thing worth being clear on early: an LLM is trained on a fixed dataset up to a certain point. It isn't pulling live answers from an infinite, ever-updating pool. Once training is done, it generates responses based on the patterns it picked up during that training, not by browsing the internet in real time.

What Happens When You Send a Message to ChatGPT?

Say you type "Hi Cutie 🙂" and hit send. It feels instant, but a few distinct steps happen in that gap.

Processing your message. The moment you hit send, your text gets converted into a form the model can actually work with (more on that in the tokenization section below). It doesn't stay as plain English for long.

Generating a response. This is the part people get wrong most often. The model isn't looking up your question in some giant database and returning a matching answer. It's predicting, one token at a time, what the most likely next piece of text should be, given everything that came before it. It does this over and over, token by token, until it decides the response is complete.

Tokenization

A token is a chunk of text. Depending on the tokenizer, that chunk could be a whole word, part of a word, or in some cases a single character.

Tokenization exists because a model can't realistically hold a fixed entry for every possible word in every language, including words people haven't invented yet. Breaking text into smaller reusable pieces lets the model handle rare words, typos, and made up terms by combining fragments it already knows, instead of failing on anything outside a fixed dictionary.

Words and tokens are not the same thing, and that trips people up constantly. A short common word like "the" is usually a single token. But something like "unbelievable" might get split into pieces like "un", "believ", and "able". Even a name like "ChatGPT" can get broken into several tokens rather than staying as one unit.

Once the text is broken into tokens, each token gets converted into a vector embedding, a list of numbers that represents that token's meaning. This is the step that actually gives the numbers significance. Words with similar meaning end up positioned close to each other in that numerical space, which is what lets the model reason about meaning instead of just symbols.

Transformers

The Transformer is the architecture almost every modern LLM is built on. It was introduced in a 2017 paper, and it's genuinely the reason the field moved as fast as it did afterward.

Before Transformers, models processed text sequentially, one word after another, using architectures like RNNs. That's slow, and it makes it hard for the model to connect something at the start of a long passage with something much later in it. Transformers process the whole sequence at once and use a mechanism called self-attention to directly relate every word to every other word, regardless of how far apart they are.

Two ideas inside the Transformer do most of the heavy lifting for understanding language:

Positional encoding. Because Transformers process everything in parallel instead of one word at a time, they lose the natural sense of order that a sentence has. Positional encoding adds information about where each token sits in the sequence, so "dog bites man" and "man bites dog" don't get treated as the same thing just because they contain identical words.

Self-attention. This is what lets the model figure out meaning from context rather than from the word alone. Take "Ruchi learned Python in college" versus "The python slithered through the grass." Same word, completely different meaning. Self-attention lets the model weigh every other word in the sentence when interpreting one word, so it picks up on "learned" and "college" to know the first is a programming language, and "slithered" and "grass" to know the second is a snake.

That combination, parallel processing plus attention across the full sequence, is what made it possible to train on massive datasets efficiently while still capturing long range context. That's exactly what language needs to be understood properly. GPT, Claude, Gemini, Llama, nearly every major model today is a variant of this same architecture, just at different scales with different training choices layered on top.

The Finishing Touches: Softmax, Temperature, and Add & Norm

There are a few more pieces working quietly in the background that decide how the model actually picks its next word, and how it stays stable while doing all this across dozens of stacked layers.

Softmax. The model's raw output is just a score for every possible next token, numbers, not probabilities. Softmax converts that list into an actual probability distribution that adds up to 100%, so instead of "token A scored higher than token B," you get "token A: 62% chance, token B: 20%." That's what turns raw scoring into something the model can actually sample a next token from.

Temperature. The dial for how safe or adventurous that sampling is. Low temperature sticks close to the highest-probability token, predictable, but can feel repetitive. High temperature flattens the distribution so lower-ranked tokens get a real shot too, more variety, less consistency. That's why the same prompt at high temperature gives you differently worded answers each time, while low temperature gives close to the same one. Neither setting is "correct," it depends on whether you want reliability or variety.

Add & Norm. Right after self-attention, and again after the feed-forward step, every Transformer block does two things. "Add" is a residual connection: the block's input gets added back to its output, like a relay runner who's allowed to tweak the baton slightly but must hand back something recognizable, not a different object. This keeps information from diluting as it passes through layer after layer. "Norm" rescales the numbers afterward to keep them in a stable range, without it, stacking dozens of these blocks makes training unstable or makes it stop working entirely. Less a nice-to-have, more a load-bearing wall.

Transformer Visualizer

If you want to see this flow instead of just reading about it, I built a small visualizer that walks through tokenization, embeddings, and next-token prediction step by step: transformer-visualizer-omega.vercel.app.

It uses a GPT-style decoder-only setup with live next-token prediction, so you can actually watch a prompt turn into tokens and then into a generated response instead of just imagining it.

Summary

  • An LLM is trained on a fixed dataset and generates replies by predicting the next token repeatedly, not by looking up or copying from the internet.

  • Before any of that prediction can happen, your text has to be broken into tokens and converted into numerical vector embeddings, since computers only ever work with numbers.

  • The Transformer architecture processes an entire sequence in parallel and uses self-attention so every word can directly relate to every other word, which is why it replaced older sequential models like RNNs.

  • Positional encoding keeps word order intact despite parallel processing, and Add & Norm keeps information stable as it passes through dozens of stacked layers.

  • Softmax turns the model's raw scores into real probabilities, and temperature decides whether the model plays it safe with the top choice or takes more risks with lower-probability tokens.

Resources where I learned from :

Originally published on Hashnode