Part 1 · The Model Under the Product

Recap (Part A) · What LLMs Are + Hallucinations

Training essence / Token / Base→SFT→Chat / four hallucination types and root causes

THE QUESTION THIS PAGE ANSWERS

ANSWER FIRST

What is the key idea behind “Recap (Part A) · What LLMs Are + Hallucinations”?

Training essence / Token / Base→SFT→Chat / four hallucination types and root causes

DECISION RULE

Make the claim earn its place. Use this page as a decision aid, not a definition to memorize. Connect the idea to one real task, one observable result, and one failure that would change your mind.

TRY NEXT

Write one question you could answer with evidence after trying this idea.

WATCH FOR

A conclusion that sounds complete but leaves the key assumption untested.

1 · What is an LLM?
🧠
What is an LLM?
Three core mental models you must have
Training Essence
LLM = A Massive Probabilistic Prediction Machine
Training = repeatedly predicting the next Token on vast amounts of text. The parameters memorize statistical patterns, not understanding. At inference, parameters are frozen — only conditional probability calculations occur.
P( next Token | all known Tokens )
Token & Context Window
Token ≠ Character · Context Window is Your "Desktop"
Each Chinese character is ~1 Token (common words cost less, rare characters more); each English word is ~1 Token. The Context Window determines how much the model can see. Anything beyond it is truncated — gone forever.

Current leading windows: Qwen 3.6 (1M), Kimi K2.5 (200K), Claude 4.6 (200K), GPT-5.4 (128K)
Evolution Path
Three Stages: Base → SFT → Chat
Base Model (text continuation) → SFT instruction fine-tuning (learns conversation format) → Chat API (simulates multi-turn dialogue).

Every API call you make is essentially constructing a carefully designed Message List, so the model continues it into exactly what you want.
[system] + [user/assistant …] + [user_now]
2 · Hallucination: LLM's Innate Limitation
👻
Hallucination: LLM's Innate Limitation
Cannot be eliminated — only mitigated
Factual Hallucination
Fabricating non-existent facts, data, or citations
Source Hallucination
Citing papers, links, or authors that don't exist
Reasoning Hallucination
Correct premise but flawed reasoning steps
Code Hallucination
Calling APIs or functions that don't exist
Root Cause 1
Flawed Parametric Knowledge
The training data contained misinformation; the model knows nothing beyond its knowledge cutoff date.
Root Cause 2
Context Misinterpretation
Unclear Prompt forces the model to guess intent; when context is contradictory, it always picks the most probable continuation — most probable ≠ most accurate.
Key Insight
→ Hallucination cannot be fully eliminated — it is an inevitable product of probabilistic prediction
→ After PreTraining, parameters are frozen and cannot automatically update knowledge — this is the fundamental cause of hallucination
→ The right approach: use engineering techniques to mitigate it — don't expect hallucinations to disappear as models get smarter

How “1 · What is an LLM” changes an answer

“Training essence / Token / Base→SFT→Chat / four hallucination types and root causes” shows that a model does not process the “word count” we see. It processes Token pieces. Tokenization affects input length, how much context fits, and how much computation a request consumes.

Length, information, and context are different

As “Training essence / Token / Base→SFT→Chat / four hallucination types and root causes” grows, separate three questions: how many Tokens the text becomes, which pieces can change the current decision, and whether older material has fallen outside the context window. Removing repetition is often more useful than simply making the window larger.

Keep what can change the decision

Use “Training essence / Token / Base→SFT→Chat / four hallucination types and root causes” as an A/B test: keep the same question while removing repeated background, compressing format, and trimming irrelevant history. Compare answer quality, latency, and Token count.

From “1 · What is an LLM” to “2 · Hallucination: LLM's Innate Limitation”

“1 · What is an LLM” grounds the problem in “🧠 What is an LLM? Three core mental models you must have Training Essence LLM = A Massive Probabilistic Prediction Machine Training = repeatedly predicting the next Token on vast amounts of text . The paramete…”. “2 · Hallucination: LLM's Innate Limitation” then moves it toward “👻 Hallucination: LLM's Innate Limitation Cannot be eliminated — only mitigated Factual Hallucination Fabricating non-existent facts, data, or citations Source Hallucination Citing papers, links, or authors tha…”. Together, they show that the lesson is not just a conclusion to remember, but a claim with conditions.

Carry the judgment into the next situation

For long text, keep what can change the conclusion before compressing format and history. A larger context is worth its cost only when the added information is useful.

  • “1 · What is an LLM”: 🧠 What is an LLM? Three core mental models you must have Training Essence LLM = A Massive Probabilistic Prediction Machine Training = repeatedly predicting the next Token on vast amounts of text . The paramete…
  • “2 · Hallucination: LLM's Innate Limitation”: 👻 Hallucination: LLM's Innate Limitation Cannot be eliminated — only mitigated Factual Hallucination Fabricating non-existent facts, data, or citations Source Hallucination Citing papers, links, or authors tha…

The final “Finish by testing the claim” brings the discussion to “Training essence / Token / Base→SFT→Chat / four hallucination types and root causes”. The useful thing to carry forward is knowing which judgments must be revisited when input, scale, or risk changes.

Mark as learned Your reading progress updates automatically
← PreviousNext →

Keep reading

The next useful article in the thread.

ARTICLE DISCUSSION

Leave one useful thought here.

Keep the idea that clicked, the question that stayed open, or a small note for the next learner.

Discussing Recap (Part A) · What LLMs Are + Hallucinations The Model Under the Product
3discussionsArticle discussion · synced with the Circle
View in the learning circle
AM
Asha MorganContent editor
INSIGHTField note

I turned one judgment from this article into a small experiment I could run today. Knowing what to observe next is more useful than simply remembering the conclusion.

ARTICLE DISCUSSION7 helpful
LH
Lin HarperIndie developer
INSIGHTInsight

After reading this, I first looked for the conditions behind the idea instead of copying the method into a project. That order made the later trade-offs much clearer.

ARTICLE DISCUSSION5 helpful
KM
Kiki MooreProduct operations
QUESTIONQuestion

When this judgment reaches real work, which constraint should be added first? I am curious which step matters most between reading and the first practical attempt.

ARTICLE DISCUSSION4 helpful