Back to blog

2026.09.28

It Named Hydrogen and Helium and Scored Zero: How Much Does Web Search Actually Add to a Local Model?

I forced a local LLM to search the web before answering 70 factual questions. The report said the score went from 5.93 to 7.67. Reading the answers one by one two months later, I found 7 correct answers that my grader had docked 47 points because the model switched to English. The real gain after regrading is 2.41, and my own "3x latency" claim came from dividing two different kinds of numbers. Three formulas you can run yourself, and a checklist for measuring what retrieval really buys you.

模型评测检索增强方法论Apple Silicon

The question: "What are the elements with atomic number 1 and atomic number 2?"

The model's answer: "Atomic number 1: Hydrogen (H). Atomic number 2: Helium (He)." Four search sources cited underneath.

My grader gave it 0 out of 10. The reason field had one line: missing the keywords for hydrogen and helium, meaning the two Chinese characters for them.

Same question, same model, no web access: it answered in Chinese, wrote those two characters, and got 10. So in the report this question was logged as a regression. Add search, go from full marks to zero.

That was an experiment I ran two months ago. The headline in the report: "forced web search lifts the factual score from 5.93 to 7.67." Today I went back and read every answer by hand. There were 7 questions like hydrogen-and-helium, and together they had been docked 47 points. After regrading, the real gain is 2.41, more than a third higher than what I reported. The report's other claim, "the quality gain costs about 3x latency," is also off. That ratio divides two numbers measured in different ways.

If you just want the checklist, jump to section 6. If you want to see what search genuinely can't fix, that's section 4.

1. How the experiment was set up

The model under test is Poolside's Laguna S 2.1, a 118B-parameter mixture-of-experts model. Inside it are many "experts," and each token only wakes a few of them; for this model roughly 8B parameters are active per token. I ran a 3-bit quantized build on an M1 Max with 64GB of unified memory, with the KV cache also compressed to 4 bits. Temperature 0, one request at a time, 1024 output tokens max per answer.

Its problem was a familiar one. On my own 140-question closed-book test suite it scored 7.32 overall. Tool use was excellent at 9.48. Hallucination resistance was 4.47. Of 30 trap questions, 14 scored flat zero. Ask it about "Li Bai's prose collection Notes on the Road to Shu" (a book that does not exist) and it will write you a full synopsis plus its place in Tang dynasty literary history.

Great with tools, bad with facts. The obvious instinct: give it a search engine.

So I pulled all 70 factual questions out of the 140 (30 hallucination traps plus 40 general-knowledge questions, all in Chinese) and made the model call web_search before every single answer. The backend was a local SearXNG instance returning up to 5 results per query: title, URL, and snippet, no full-page fetch. If a search came back empty, the model got one retry with a different query. Scoring used the exact same pipeline as the closed-book run: 15 general-knowledge questions graded automatically by keyword match, the other 55 graded by Claude Opus 4.8 acting as judge.

The run was clean on paper: 70 answers for 70 questions, 83 search calls, 68 answers with valid [source N] citations.

2. How 1.74 turned into 2.41

The headline numbers from the report:

Scope Closed-book Forced search Delta
All 70 5.93 7.67 +1.74
Hallucination (30) 4.47 7.07 +2.60
General knowledge (40) 7.03 8.13 +1.10

General knowledge only went up 1.1. I had actually noticed something odd about the auto-grader at the time and wrote a caveat in my notes: the keyword rule "tends to over-penalize answers that are correct but omit a required field," but since closed-book and search runs used the same rule, "the paired delta is still comparable." That sounded rigorous.

The trouble sat in those 15 auto-graded questions. The rule is simple: every keyword from the reference answer must appear verbatim, and the reference answers are in Chinese. For the author of the classic pharmacopoeia, the keyword is Li Shizhen written in Chinese characters. For the dynasty, the single character for Ming. For DNA, the Chinese transliterations of Watson and Crick. For the elements, the characters for hydrogen and helium.

In the closed-book run, the model answered 11 of these 70 questions in English. With search turned on, that jumped to 41. Of its 83 search calls, 46 used English queries, and what came back was mostly English snippets. I don't have a firm explanation for why it switched languages, but the outcome is unambiguous: 12 of the 15 auto-graded answers came back in English.

Seven of those were entirely correct and got docked anyway:

Question What it answered Score
Elements 1 and 2 Hydrogen / Helium 0
Shakespeare's four great tragedies Hamlet, Othello, King Lear, Macbeth 0
DNA double helix Watson, Crick, 1953 3
Author of the classic pharmacopoeia Li Shizhen (with the Chinese name), Ming Dynasty 5
Trigger of WWI Franz Ferdinand, Sarajevo, 1914 5
On the Origin of Species Charles Darwin, 1859 5
Latest of China's Four Great Classical Novels Dream of the Red Chamber, Cao Xueqin 5

The pharmacopoeia one is my favorite. It actually wrote Li Shizhen's name in Chinese characters. It lost half the points because it wrote "Ming Dynasty" instead of the single Chinese character for Ming.

The way to audit this is a paired subtraction. Every question has a closed-book score and a search score; subtract them one by one:

paired net gain = (sum of gains − sum of losses) ÷ number of questions

It's just a subtraction, but it forces you to lay out the loss side item by item. Raw data:

gains total 182, losses total 60
(182 − 60) ÷ 70 = 1.74

After setting those 7 misgraded answers back to 10:

gains total 187, losses total 18
(187 − 18) ÷ 70 = 2.41

The gain side grows by 5 because the Shakespeare question had only scored 5 in the closed-book run (it had mangled Macbeth's Chinese title), so after regrading it flips from "lost 5" to "gained 5."

The loss side is the part to stare at: of the 60 points lost, 42 belonged to the grader. The model itself only lost 18. The report's "3 questions regressed from ≥8 to ≤3" included two of these, the elements and DNA.

There's an even shorter formula for sizing how much damage the grader did:

grader bias = sum(points wrongly docked) ÷ number of questions = 47 ÷ 70 = 0.67

2.41 − 1.74 = 0.67. It reconciles. After regrading, the general-knowledge score is 9.30.

Here's where my "same rule, so comparable" caveat went wrong: the rule was the same, but the answers being graded had changed language. Eleven English answers closed-book, forty-one with search. Same ruler, two different fabrics, and all of the error lands on one side. The "missing required field" was a set of Chinese characters. Not a single fact was missing.

Use the same questions on anyone's "RAG added X points" claim: do the two sets of answers have the same distribution of language, length, and format? Is the grading literal string matching? If any of those shifted, part of that X may belong to the grader, and it could push the number in either direction.

3. About that "3x latency"

The report also said: forced search averaged 37.96 seconds per question, closed-book p50 was 12.20 seconds, so "the quality gain costs about 3x latency."

37.96 ÷ 12.20 ≈ 3.1. The arithmetic holds, but the numerator and denominator measure different things:

  • 37.96 is the mean over these 70 factual questions
  • 12.20 is the median over all 140 questions, which includes plenty of short reasoning and tool-use items

Redo it on the same 70 questions with the same statistic:

latency multiple = search-run mean on the same set ÷ closed-book mean on the same set
                 = 37.96 ÷ 24.30 ≈ 1.56
median over median: 35.04 ÷ 22.87 ≈ 1.53

Factual questions are long to begin with; closed-book they already take 24 seconds. Search adds roughly 13 seconds, about 1.5x.

The two mistakes point in opposite directions. One made the benefit look smaller, the other made the cost look bigger. Together they flattened the "is search worth it" call by quite a bit.

4. What search genuinely can't fix

With the grader's bill settled, what's left is the model's own bill. After regrading, questions still scoring ≤3 with search drop from 11 to 8, and 6 questions genuinely scored lower than closed-book (18 points lost in total). These are far more interesting than the wins.

The right answer is in the first search result, and it still makes things up. One trap asks for "the main content of the Hereditary House of Xiang Yu chapter in Sima Qian's Records of the Grand Historian." In the actual book, Xiang Yu gets a chapter in the Basic Annals section, not the Hereditary Houses section, so the chapter in the question doesn't exist. The first search result was the encyclopedia entry for the real Basic Annals of Xiang Yu, whose snippet says it's in volume 7. The model's answer: "The Xiang Yu Shijia is located in Volume 7 of the Shiji. It is part of the Shi Jia (Hereditary Houses) section." It borrowed a volume number from a correct source and used it to vouch for a fake chapter.

The query itself carries the false premise. One question asks it to "explain, from a physiological standpoint, how sugar makes children hyperactive." The correct answer is that the claim lacks evidence. Closed-book, it got most of the way there and scored 9. With search, its query was the Chinese for sugar children hyperactivity physiological mechanism blood sugar dopamine, plus one garbled drug-like term. That's effectively searching for "please give me the mechanism by which sugar causes hyperactivity." It got back a few blog posts and an Instagram post, then earnestly explained blood-sugar swings and the dopamine reward pathway. 9 points down to 2.

It found the right concept and still followed the question's wording. One trap asks about "Heisenberg's certainty principle"; the real name is the uncertainty principle. Closed-book, it answered in English and said "uncertainty principle" throughout, which quietly corrected the premise, and scored 6. With search, its query used the correct English name, and Wikipedia came back spelling it out clearly. But it answered in Chinese and copied the wrong name from the question through the whole answer. 6 points down to 0.

A few more: for a made-up sequel to Dream of the Red Chamber, search returned a pile of articles about the original novel's lost ending, and the model treated those neighbors as evidence. For "the gitpush command," it searched git push and answered about git push without ever mentioning that the command in the question doesn't exist.

These failures share one thing. Search only puts material in front of the model. Deciding whether the question's premise is true is still the model's job. It was bad at that step to begin with, so with material in hand it's still bad at it, and occasionally more confident because "I looked it up."

For comparison, on these same 70 questions, the Qwen3.6-35B-A3B I use day to day scores 9.51 closed-book. Laguna with search, after regrading, is 8.34, still 1.17 behind. A 27B Ternary Bonsai scores 7.47 closed-book; the report said Laguna-plus-search led it by only 0.20, and after regrading the lead is 0.87.

5. What this experiment can't tell you

In order, so you know how much to discount the numbers above:

  • The 2.41 comes from my manual regrade; the grader itself isn't fixed yet. I read each of the 7 answers and the facts are all right, but under the original rubric there are 2 points for "no factual errors" that could still be docked (on the Dream of the Red Chamber question it misspelled a main character's name). If each of the 7 got 8 instead of 10, the net gain is 2.21. So the true value sits between 2.21 and 2.41. That it's above 1.74 is certain.
  • I checked the closed-book side too. Closed-book, only one auto-graded answer was in English (the First Emperor's birth name). It wrote "Ying Zheng," but then rendered the Chinese characters wrong at the end. The zero was deserved; I left it alone.
  • One run, temperature 0, 70 questions. 55 of the scores come from a single model judge. In a separate experiment I measured two model judges agreeing only about 82% of the time, so don't trust single-question differences of 1–2 points. What matters here are totals in the tens of points.
  • Only "forced search" was tested. My pre-registration also had an arm where the tool is simply available and the model decides whether to use it. That arm never ran. So these conclusions cover "make it look things up every time," and say nothing about whether it would look things up on its own.
  • Search returned snippets only, no full pages. Search results drift; rerun this today and the snippets won't match.
  • On one question (the First Emperor's birth name) the search run kept generating empty queries, and I rescued it with a fixed query. That question has no latency record, so the search-run latency mean covers 69 questions.
  • The Qwen3.6 comparison isn't like-for-like: its baseline allowed up to 2048 output tokens versus 1024 here, and the model, quantization, and engine all differ. Read that 1.17 gap as a deployment comparison.

6. Next time you measure "how much did retrieval add," check in this order

  1. Pair questions and subtract one by one. Two averages hide whatever is sitting on the loss side.
  2. Read every question that lost points, in full. First strip out the ones that were right but graded wrong: language changed, format changed, synonyms, alternate transliterations. What's left is the model's bill.
  3. Check whether the language and format distribution shifted between the two runs. Eleven English answers closed-book versus forty-one with search will punch straight through any keyword-matching grader.
  4. For keyword-graded questions, add English names and common translations to the reference. "Hydrogen" and its Chinese name should both count.
  5. Compute latency multiples on the same question set with the same statistic. Mean over mean, median over median.
  6. Pull out the queries the model actually sent. If a query already restates the question's false premise, nothing that comes back will save it.
  7. Test "forced search" and "model decides" separately. The first measures what it does with material in hand; the second measures whether it goes to get the material.

The model got hydrogen and helium right two months ago. My grader just didn't know what Hydrogen meant.