An Introduction to the Concepts and Logic Behind Convolutional Neural Networks and Transformer-Based Generative AI
This post was translated from Chinese by AI. If anything reads oddly, the Chinese original is authoritative. 中文原文
Notes, more or less
Basic Concepts
Convolutional Neural Networks (CNN): AI's “Eyes” 👀
Convolutional Neural Networks (Convolutional Neural Networks, CNN) are a core technology that gives computers “vision.” They excel at processing grid-like data such as images.
- How they work: Imagine holding a small square frame (called a “convolution kernel” or “filter”) and sliding it across a large photo. At each position, you focus only on what's inside the frame, looking for specific features such as edges, lines, or patches of color.
- Layer by layer: CNNs have many layers. The first layer might recognize only simple lines; the second combines those lines into shapes (such as circles and squares); deeper layers can recognize complex objects (such as a cat's ears or a car's wheels).
Generative AI (Generative AI): AI's “Creativity” 🎨
Generative AI doesn't just analyze existing data; its goal is to create entirely new data.
- How it works: These AI models have read vast numbers of books and viewed countless artworks, learning the “patterns” or “probability distributions” behind that data. Once they learn these patterns, they can predict and generate new text, images, or even videos based on your prompt (Prompt).
- More than copying: Rather than searching a database for ready-made answers, they work like human artists, using what they've learned to build entirely new content, stroke by stroke (or word by word).
How They Differ and Work Together
The Core Difference: How Do They Process Information?
-
CNN (Convolutional Neural Networks) —— Focus on “Local Features” 🔍
- Logic: Their strength is feature extraction. As mentioned earlier, they scan an image piece by piece with a magnifying glass. They assume “neighbors matter” (for example, pixels representing a nose are usually near pixels representing eyes).
- Limitations: Because they focus so heavily on local features, they sometimes overlook relationships across the whole image. For example, if you swap the mouth and eyes in a photo of a face, a traditional CNN might still identify it as “a face”: it detects eyes and a mouth without paying much attention to whether they're in the right places.
-
Large Generative Models (Based on Transformer) —— Focus on “Global Relationships” 🧠
- Logic: Modern generative AI (such as GPT, Gemini) uses the Transformer architecture. Its key technique is the attention mechanism (Self-Attention).
- Capabilities: It is no longer limited to “neighbors.” It can attend to all the input information at once and calculate long-range relationships between different parts.
- Example: When processing the sentence “Apples taste good because this fruit is sweet,” it can instantly connect “apples” at the beginning with “fruit” later on, even when many words separate them.
Multimodality: How Do They Work Together? (Topic 3)
Today's multimodal models (AI that can describe images) often combine the two.
- Step One (Eyes): The AI may use a CNN (or a similar vision encoder) to “look at” an image and extract key visual features (lines, shapes, objects).
- Step Two (Brain): These features are converted into data and fed to a Transformer (a large language model). The model combines this visual information with its language capabilities to generate a description: “This cat is sitting on the sofa.”
Transformer Architecture: Background and Significance
Transformer is a deep learning model introduced by Google in 2017. Before this, AI processed language as if reading a running log, often forgetting earlier parts as it moved along. Transformer can “see” the entire sentence at once and understand how each word relates to every other word, regardless of how far apart they are.
Transformer Architecture: Mechanisms and Structure
Core Mechanism: Self-Attention (Self-Attention)
In the sentence “The animal didn't cross the street because it was too tired”:
- When the model reads the word “it”, the self-attention mechanism calculates its relevance to every other word in the sentence.
- Because it sees “tired” later on (which usually describes living things), it assigns a very high score to “animal” and a low score to “street.”
- This is “understanding context”: rather than translating mechanically, the model uses this mathematical “scoring” to work out what words actually refer to.
Overall Structure: Encoder and Decoder (Encoder & Decoder) 🏗️
Transformer was originally designed for translation tasks (such as translating English into French). Its structure resembles a super translator, split into two halves:
-
Encoder (Encoder) —— “Responsible for Reading Comprehension” 📖
- Role: It breaks down the input (such as an English sentence), understands its grammar, meaning, and context, then compresses it into a set of numbers (vectors).
- Analogy: When listening to someone, you don't just hear sounds—you distill their meaning in your mind.
-
Decoder (Decoder) —— “Responsible for Generating Output” ✍️
- Role: It takes the “meaning” supplied by the encoder and starts generating output word by word (such as French text or the next paragraph). With each word it generates, it looks back at the previous output to keep things coherent.
- Analogy: Once you understand what someone means, you put your response into words.
How the Transformer Architecture Processes Data
Let's break the whole process into three core stages: Data input (turning text into numbers) -> Internal model processing (mathematical calculations) -> Training loop (correcting errors).
This is what actually happens inside a GPU every second:
Stage One: Data Input —— From Text to “Vectors” 🔢
Computers don't understand words; they only understand numbers. So the first step is translation.
-
Tokenization (Splitting Text into Tokens):
- Input sentence: “I like AI.”
- The model splits it into small chunks (Token):
["我", "喜欢", "AI"]. - It looks them up in a dictionary (vocabulary) and converts them into IDs:
[105, 2044, 89].
-
Embedding (Embedding) —— This Is the Crucial Step! 💎
- Simply converting “I” into the number
105isn't enough, because105and106have no logical relationship. - The model converts each ID into a long vector (Vector). It's like giving each word an “ID card” containing hundreds of feature values.
- Data transformation: For example, the word “king” might become
[0.9, 0.1, -0.5, ...](representing: male, powerful, human...). - The remarkable part: In this mathematical space, king - man + woman ≈ queen. Words with similar meanings are also close together in this space.
- Simply converting “I” into the number
-
Positional Encoding (Position Encoding):
- Because Transformer looks at all words at once (rather than reading sequentially like humans), it doesn't know the difference between “I hit you” and “You hit me.”
- So we need to add a number representing position to each word's vector, telling the model which word comes first and which comes later.
Stage Two: Inside the Model —— The Famous Q, K, V Matrix Operations 🧠
The data is now a set of matrices (tables of numbers), which enter Transformer's core component: the self-attention layer (Self-Attention).
What happens here? The model performs a massive “database query.” For each word in the sentence, it generates three vectors:
- Query (Q - Query): What am I looking for?
- Key (K - Key): What features do I have? (Like the labels on book spines in a library)
- Value (V - Value): What is my actual content?
How the data flows:
-
Calculating Relevance (Attention Score):
- The model takes word A's
Qand computes its dot product (multiplication) with every other word'sK. - A large result means the two words are closely related.
- Example: Multiplying the
Qof “it” by theKof “cat” gives a high score; multiplying it by theKof “table” gives a low score.
- The model takes word A's
-
Weighted Sum:
- Using those scores, the model adds together the
Vvectors (content) of all the words to form a new vector. - Result: The vector that originally represented “it” now incorporates information about “cat,” making it richer.
- Using those scores, the model adds together the
Stage Three: The Training Loop —— How Does It Learn? (Backpropagation) 📉
This is the part you're most interested in: What happens at each training step?
It's a continuous process of “guess -> get penalized -> correct”.
Scenario: We feed a sentence to the model for training —— “The sky is [blue].” We hide “blue” and let the model guess.
Step 1: Forward Pass (Forward Pass) —— A Wild Guess
- The data goes through hundreds of millions of matrix multiplications (the QKV calculations mentioned above), and the model finally outputs a list of probabilities.
- Model output: It might assign a 60% probability to “green” and a 10% probability to “blue.”
- Correct answer: “Blue” should have a probability of 100%.
Step 2: Calculating Loss (Loss Calculation) —— Tallying the Error
- We use a mathematical formula (usually the cross-entropy loss function Cross-Entropy Loss) to measure how far off the model is.
- Since the model guesses “green” but also assigns some probability to “blue,” the error might be a number like
Loss = 2.5. If it correctly guesses “blue,”Losswill be close to0.
Step 3: Backpropagation (Backward Pass) —— Finding Who's to Blame 🔥
- This is the most remarkable step in AI training.
- Now that we've calculated
Loss = 2.5, we need to know: what caused this error? - Using calculus (the chain rule), the computer works backward from the output to calculate how much each layer and each parameter (weight) “contributed” to the error. This is the gradient (Gradient).
- In plain terms: The system discovers, “Oh, that parameter in the matrix in layer 3 is too large, causing it to favor ‘green.’”
Step 4: Updating Parameters (Optimizer Step) —— Making Corrections
- Once we know which parameter is responsible, the optimizer (Optimizer) (such as Adam) adjusts it slightly.
- Action: Make it a tiny bit smaller.
- Goal: The next time the model sees “The sky is...,” its probability of guessing “blue” will be a little higher.
Summary: What Exactly Happens to the Data?
- Input: Text -> table lookup -> vectors.
- Computation: Vectors flow through hundreds of network layers, continually mixing contextual information (QKV operations).
- Output: A probability distribution (tens of thousands of words in the vocabulary, each with a probability of appearing).
- Error correction: Calculate the error -> calculate the gradients -> update all parameters.
When training large models like GPT-4, this process repeats trillions of times.
6 Questions About Data Processing in the Transformer Architecture
1. "Embedding" Explained: Turning a Dictionary into a “Map” 🗺️
The core goal of Embedding (embedding) is to turn discrete IDs (integers) into continuous vectors (coordinates).
-
Why can't we just use IDs?
- Suppose you use IDs:
苹果(10),香蕉(11),电视(500). - The computer sees that
10and11are close, so apples and bananas are similar (which is fine). - But mathematically,
500is 50 times10. Does that mean a TV is 50 times an apple? That clearly makes no sense. IDs are just labels, with no semantic relationships.
- Suppose you use IDs:
-
How does Embedding work?
- We assign each word coordinates in a high-dimensional space (for example, 768 dimensions, or 12288 dimensions in GPT-3).
- Imagine a vast multidimensional space. At the start of training, these coordinates are randomly generated.
- As training progresses, the model discovers that “apple” and “banana” often appear in similar contexts (such as “eat” and “sweet”). Through backpropagation, it pulls their coordinates closer and closer together.
-
What is the result?
-
The Embedding layer is like a huge lookup table (Lookup Table).
-
Input
ID: 10-> output vector[0.1, -0.5, 0.9, ...]. -
In this space, mathematical operations take on semantic meaning:
$$ ec{v}(\text{King}) - \vec{v}(\text{Man}) + \vec{v}(\text{Woman}) \approx \vec{v}(\text{Queen}) $$
-
2. How Q, K, and V Actually Work: Vector Dot Product (Dot Product) 📐
You've hit on the key question: what exactly does multiplying Q by K do?
Here, “multiplication” specifically means the vector dot product (Dot Product).
-
Mathematical formula: $\text{Score} = Q \cdot K^T$
-
Geometric meaning: The dot product is an excellent tool for measuring the similarity of two vectors (whether they point in the same direction).
- If vectors A and B point in the same direction, their dot product is largest (positive).
- If they are perpendicular (unrelated), their dot product is 0.
- If they point in opposite directions, their dot product is negative.
-
The process:
- Query (Query): “I want to find words related to the word ‘I.’”
- Key (Key): “I'm ‘cat,’ and here are my feature labels.”
- Dot product: $Q_{\text{I}} \cdot K_{\text{cat}}$. If these two vectors point in the same direction (giving a large value), “I” and “cat” are closely related in the current context.
- Softmax (normalization): After calculating all the dot products, we use the Softmax function to turn them into probabilities (which sum to 1). For example,
[0.8, 0.1, 0.1]. This means the model decides to put 80% of its attention on “cat.” - Weighted sum: Finally, use these probabilities to weight V (Value).
Summary: $Q \times K$ is essentially similarity matching, determining where the model should “focus.”
3. Forward Pass: Does It Predict the Next Token Based on the Previous Tokens?🔮
Yes, that's exactly how generative models like GPT (Decoder-only) work.
-
Autoregression (Autoregressive):
- Input:
[A, B, C] - Forward Pass computation...
- Predicted output:
D(the one with the highest probability). - Next step: Append
Dto get[A, B, C, D], then run another Forward Pass to predictE.
- Input:
-
“Cheating” during training (Teacher Forcing):
- For efficiency, we don't actually generate one character at a time during training. We feed the entire sentence
[A, B, C, D, E]into the model at once. - But! We use a mask matrix (Mask) to hide the future.
- When the model sees
B, the Mask hidesC, D, E, forcing it to use onlyAandBto predictC. This keeps the training logic consistent with the inference logic.
- For efficiency, we don't actually generate one character at a time during training. We feed the entire sentence
4. Cross-Entropy Loss (Cross-Entropy Loss): How Is the Score Calculated?📉
This is a measure of “how accurate the guess is.”
-
Scenario: The model predicts that the next word is “cat.”
- The vocabulary contains 10000 words.
- Correct label (Ground Truth):
ID: 5(cat). Mathematically, this is a One-hot vector:[0, 0, 0, 0, 0, 1, 0, ...](only position 5 is 1; all others are 0). - Model prediction (Prediction): After Softmax, the model outputs a probability distribution:
[0.01, 0.02, ..., 0.8, ...](it assigns 80% to cat and 20% to other words).
-
Formula:
$$ ss = - \sum (y_{true} \cdot \log(y_{predicted})) $$
-
Calculation:
- Since only the “cat” position in the true label is 1 and all others are 0, the formula simplifies to: look only at the predicted probability for “cat.”
- $$ Loss = - \log(0.8) $$
- $\log(0.8) \approx -0.223$
- $Loss = -(-0.223) = 0.223$
-
Intuition:
- If the model assigns “cat” a probability of 1.0 (100% confident), $-\log(1) = 0$. Loss = 0 (perfect, no penalty).
- If the model assigns “cat” a probability of 0.1 (not very confident), $-\log(0.1) = 2.3$. Loss = 2.3 (a large penalty).
That's how Loss is calculated: the lower the probability, the greater the penalty (Loss).
5. The Chain Rule and Layers: Passing Messages Through a High-Rise 🏢
-
The concept of layers (Layers):
- A Transformer doesn't have just one Attention layer; it stacks many layers (like a burger).
- GPT-3 has 96 layers. BERT-Large has 24 layers.
- Data is processed by the first layer, whose output is passed to the second, and so on. Each layer has its own parameters (weight matrices).
-
How the chain rule (Chain Rule) applies here:
-
Once we've calculated the Loss (say, 0.223), we need to assign “blame” to every parameter across the 96 layers.
-
Mathematical logic: The “chain” here is the nesting of functions between layers.
$$ = f_3(f_2(f_1(x))) $$
-
To find the derivative (gradient) of $Loss$ with respect to the first layer's parameter $w_1$, we must use the chain rule:
-
$$ rac{\partial Loss}{\partial w_1} = \frac{\partial Loss}{\partial y} \cdot \frac{\partial y}{\partial f_3} \cdot \frac{\partial f_3}{\partial f_2} \cdot \frac{\partial f_2}{\partial w_1} $$
- Intuitively: The gradient flows down from the top (Loss) like water. As it passes through each layer, it uses that layer's formula to calculate how its parameters should change. This is called Backpropagation.
6. Mask It Yourself, Work It Out Yourself: Self-Supervised Learning 🔁
Your understanding is exactly right! This is the key reason large models can get smarter by reading vast amounts of data.
-
No manual labeling needed:
- Traditional AI requires humans to label images one by one: “This is a cat,” “This is a dog.” That's too expensive and too slow.
- Transformer uses self-supervised learning.
-
The process:
- Take a book: For example, Harry Potter.
- Create a question: The model randomly masks part of a sentence: “Harry picked up the [MASK].”
- Keep the answer: The system quietly records that the answer is “wand.”
- Predict: The model guesses based on the context.
- Compare and optimize: The model guesses wrong -> calculate Loss -> update parameters through backpropagation.
- Repeat: Move to the next sentence, mask again, and guess again.
Last updated 2026-02-09
Comments 0