Part 2 · The Harness Around the Model

Streaming Output & Format Pairing

JSON needs full text to parse / MD streams char-by-char / XML renders on tag capture — live demo

THE QUESTION THIS PAGE ANSWERS

ANSWER FIRST

What is the key idea behind “Streaming Output & Format Pairing”?

JSON needs full text to parse / MD streams char-by-char / XML renders on tag capture — live demo

DECISION RULE

Make the claim earn its place. Use this page as a decision aid, not a definition to memorize. Connect the idea to one real task, one observable result, and one failure that would change your mind.

TRY NEXT

Write one question you could answer with evidence after trying this idea.

WATCH FOR

A conclusion that sounds complete but leaves the key assumption untested.

Choose a Demo Scenario
Real-time Comparison: Three Formats
JSON Waiting
Raw stream Tokens (accumulating char by char)
⏳ Waiting for the complete JSON to arrive before parsing…
JSON is malformed mid-stream,
so it cannot be parsed or rendered incrementally
✓ JSON parsing complete (after full text arrives)
Markdown Waiting
Render area (character-by-character display)
⚠️ Although character-by-character display is smooth,
you cannot reliably
extract structured fields from Markdown
XML Custom Tags Waiting
Incremental parse-render (renders as soon as </tag> is captured)
Key Takeaways
JSON Streaming
Backend-friendly but poor for streaming:
Must wait for the full text to parse,
users see a loading state throughout
Markdown Streaming
Smoothest user experience:
Character-by-character display feels like typing,
but cannot extract structured fields for programmatic use
XML Streaming
Best of both worlds: streaming + structured
Captures </name> and renders immediately,
Claude's native recommended approach
Core Takeaway: XML captures tags → renders fields immediately; JSON waits for full text → parses once; MD displays char-by-char → cannot extract fields.

How “Choose a Demo Scenario” changes an answer

“3 tools with name / type / highlight” shows that a model does not process the “word count” we see. It processes Token pieces. Tokenization affects input length, how much context fits, and how much computation a request consumes.

Length, information, and context are different

As “conclusion / reason / action / confidence” grows, separate three questions: how many Tokens the text becomes, which pieces can change the current decision, and whether older material has fallen outside the context window. Removing repetition is often more useful than simply making the window larger.

Keep what can change the decision

Use “conclusion / reason / action / confidence” as an A/B test: keep the same question while removing repeated background, compressing format, and trimming irrelevant history. Compare answer quality, latency, and Token count.

From “Choose a Demo Scenario” to “Real-time Comparison: Three Formats”

“Choose a Demo Scenario” grounds the problem in “3 tools with name / type / highlight”. “Real-time Comparison: Three Formats” then moves it toward “JSON Waiting Raw stream Tokens (accumulating char by char) ⏳ Waiting for the complete JSON to arrive before parsing… JSON is malformed mid-stream, so it cannot be parsed or rendered incrementally ✓ JSON parsing…”. Together, they show that the lesson is not just a conclusion to remember, but a claim with conditions.

Carry the judgment into the next situation

For long text, keep what can change the conclusion before compressing format and history. A larger context is worth its cost only when the added information is useful.

  • “Choose a Demo Scenario”: 3 tools with name / type / highlight
  • “Real-time Comparison: Three Formats”: JSON Waiting Raw stream Tokens (accumulating char by char) ⏳ Waiting for the complete JSON to arrive before parsing… JSON is malformed mid-stream, so it cannot be parsed or rendered incrementally ✓ JSON parsing…
  • “The closing point”: conclusion / reason / action / confidence

The final “The closing point” brings the discussion to “conclusion / reason / action / confidence”. The useful thing to carry forward is knowing which judgments must be revisited when input, scale, or risk changes.

Mark as learned Your reading progress updates automatically
← PreviousNext →

Keep reading

The next useful article in the thread.

ARTICLE DISCUSSION

Leave one useful thought here.

Keep the idea that clicked, the question that stayed open, or a small note for the next learner.

Discussing Streaming Output & Format Pairing The Harness Around the Model
3discussionsArticle discussion · synced with the Circle
View in the learning circle
AM
Asha MorganContent editor
INSIGHTField note

I turned one judgment from this article into a small experiment I could run today. Knowing what to observe next is more useful than simply remembering the conclusion.

ARTICLE DISCUSSION7 helpful
LH
Lin HarperIndie developer
INSIGHTInsight

After reading this, I first looked for the conditions behind the idea instead of copying the method into a project. That order made the later trade-offs much clearer.

ARTICLE DISCUSSION5 helpful
KM
Kiki MooreProduct operations
QUESTIONQuestion

When this judgment reaches real work, which constraint should be added first? I am curious which step matters most between reading and the first practical attempt.

ARTICLE DISCUSSION4 helpful