Back to blog

2026.10.02

Pointing directly got 186 right, parse-then-pick got 188, the hybrid got 190: which way should a local model click the screen?

I had a 35B local model click buttons in a window with no accessibility tree. Each route has a trade-off: asking for coordinates directly is over 6x faster but slightly less accurate, while parsing first is a bit more accurate but costs about 3 seconds per step. The hybrid router got 190 of 200 real clicks right, but those 2 extra wins prove nothing statistically. The more useful finding was that "both routes agree" is far less reassuring than it looks, because the two routes fail together. Three calculations you can run yourself: whether two routes fail independently, the average latency of a hybrid router, and how much a zero-error run actually tells you.

LLMAgent评测方法论

I ran two routes, 200 clicks each. One got 186 right, the other 188. Put them together and you get 190.

That looks like one plus one making more than two. I was pleased for about a minute. Then I counted the questions where the hybrid actually beat parse-then-pick: 2. Questions where it lost: 0. Two to nothing is something a coin can do.

What actually changed my mind was something else. I had assumed that when both routes agree on a target, it's safe to click. They turned out to be wrong together 7 times as often as you'd expect if their errors were independent. That's the part that matters for destructive buttons.

If you just want the check, chapter 4 is a single multiplication. If you want to know why the 190 doesn't count, see chapter 3.

1. The problem: clicking a button when the window tells you nothing

The comfortable case for a computer-use agent is an accessibility tree, where the OS hands you every button's name and position. Plenty of interfaces don't have one: games, custom-drawn canvases, Windows programs running under Proton. All you get is a screenshot.

I used the open-source cua-driver, which recently shipped an optional extension for exactly this fallback. You hand it a screenshot. A YOLOv8 icon detector (OmniParser v2, 81MB) boxes the clickable regions, and a small OCR model (PP-OCRv5 mobile, 4.8MB detector + 7.8MB recognizer) reads any text. It all runs on CPU. No large model is involved in this step, just two small detectors.

That gives two routes:

  • Parse, then pick: the extension cuts the screenshot into candidate boxes, each box gets a letter drawn on the image, and the 35B picks a letter.
  • Point directly: skip the parsing, show the 35B the screenshot, ask where the button is, and have it answer with a point.

To get ground truth I drew 25 scenes with Tk on a virtual display, deliberately leaving only a single window node in the accessibility tree. Each scene has 8 targets, half text buttons and half unlabeled icons. Tasks describe the function ("undo the last step"), never the icon name. The app logs which control each click actually landed on, and that log is the only thing that decides right or wrong.

2. Two faceplants first

First: text-only candidates, and icons collapsed. At first I passed only the parser's text and box coordinates to the model, to save an image. Text buttons went 100/100. Icons went 23 out of 96. A 4B model trained specifically for this job also got 23.

One look at the parser output explained it: every icon is labeled icon-class-0. The model gets a dozen identically named boxes and is choosing blindfolded. Each scene has 11 to 15 candidates, so a blind guess on icons lands about 9% of the time. 23/96 is 24%, so it beats guessing by only a little.

Once I sent the screenshot with the letter marks drawn on it, the 35B got 88/96 icons and the 4B got 78/96.

Second: ask for pixel coordinates, get 0/8. I wrote "return pixel coordinates" in the prompt and all 8 test questions missed. The model simply ignored that line and answered in Qwen's native 0-to-1000 relative coordinates. I treated those as pixels, so nothing landed where it should. Parsing the answer as 0–1000 fixed the same questions.

Both faceplants have the same shape. I nearly wrote down "the model can't recognize icons" and "the model can't localize." In fact the first time I'd fed it input with no information in it, and the second time the two sides disagreed about what a coordinate meant.

3. How good is each route, and how good is the hybrid

On the 196 questions both routes can attempt:

Route Correct Time per step
Parse, then pick (with image) 188 ~2.5 s
Point directly 179 ~0.37 s

Parsing alone takes about 2 seconds on CPU for a 900×600 window. Direct pointing has a p50 of 365 ms, over 6x faster, at the cost of 9 fewer correct answers. In a separate 200-question pointing run it missed 17: 6 times it said it couldn't find the target, 10 times it pointed at a different control, and once it pointed at empty space. Of the 10 wrong-control misses, 7 were icons and 3 were text buttons. All 6 "can't find it" answers were icons, and 4 of them were the same "reply to email" icon.

Then the hybrid. Ask for a point first. In parallel, crop a small patch around that point, draw a red crosshair on it, and ask "is the control under the crosshair the one you want?", reading the yes/no probability. Separately ask "is this action risky?" (deletes data, can't be undone, spends money, sends something to other people). If the verify score is ≥ 0.5 and the danger score is < 0.5, click the point. Otherwise fall back to parse-then-pick.

I tuned the thresholds on those 25 scenes, froze them, then generated 25 new scenes from fresh random seeds and made 200 real clicks:

Route Correct
Hybrid 190
Parse-then-pick only 188
Point-only 186

The two single-route rows were scored offline on the same screenshots, using each scene's control layout for hit testing. I checked that: across the 200 real hybrid clicks, the offline hit test and the app's own click log disagreed 0 times. So the offline rows can be trusted.

Now, about that 190.

When comparing two methods, ignore the totals and count the questions where one was right and the other wrong. Hybrid vs. parse-then-pick had 2 such questions, and the hybrid won both. Hybrid vs. point-only went 8 to 4.

If neither method is actually better, those discordant questions should split about evenly. You can compute the odds of them all landing on one side yourself:

k discordant questions all favoring one side: two-sided p = 2 × 0.5^k

With 2 questions, p = 0.5. You need 6 out of 6 on one side before p drops to 0.031. For the 8-to-4 split, the two-sided binomial p is about 0.39.

Put plainly, I was seeing things in that 190. On accuracy the hybrid ties parse-then-pick, and what it buys is speed. Of 200 clicks, 139 took the fast route (136 correct) and 61 fell back (54 correct). Median decision time was 0.70 s and the mean was 1.37 s. Parse-then-pick alone takes over 2 seconds per step. You can work out the average latency yourself:

mean latency ≈ fast-route share × fast-route time + fallback share × fallback time
139/200 × 0.68 + 61/200 × 2.96 ≈ 1.37 s

The measured mean was 1.37 s, so the formula checks out. The 0.68 s fast route covers three calls: point, verify, danger. The 2.96 s fallback is those three plus parsing and picking. The formula tells you where the speed comes from: the only knob is the fast-route share, and the verify and danger gates are what hold it down. Of the 61 fallbacks, the danger gate caused 33, a low verify score 22, and "can't find it" 6.

A side note on the verify gate: it's leaky. During calibration, 7 of the 11 wrong points still scored ≥ 0.5, while 150 of the 183 correct points passed. That's an 82% pass rate on correct points against 64% on wrong ones, not much of a gap. The held-out scenes had few wrong points (3 of 8 passed), which is too small a sample to say more. The gate is useful for routing. Don't count on it to stop errors.

4. How much is "both routes agree" worth

What about risky buttons, like delete or pay, where one wrong click really hurts? My rule: click only if the direct point and the parsed pick land on the same control. If they disagree, don't click, and hand it back to the human.

Scored offline on the same screenshots: of 36 risky actions, the routes agreed 31 times and were right every time. They disagreed 5 times, and those were abstentions. For comparison, parse-then-pick alone would have gotten 34 right and clicked 2 wrong. Zero errors looks great.

Then I looked across all 200 clicks:

  • The routes agreed 185 times and were wrong in 5 of them
  • Both routes were wrong 6 times in total, and in 5 of those they picked the same wrong control

You can estimate the expected number by assuming the errors are independent:

expected joint errors ≈ total × error rate A × error rate B
200 × (14/200) × (12/200) ≈ 0.84

The observed count was 6, 7 times the independent estimate.

I can roughly guess why. The two routes have different pipelines, but the same 35B makes the call in both. All 6 joint errors were icon questions: archive twice, then undo, redo, save, and stop once each. If the model misreads an icon, it keeps misreading it whichever route it takes.

So how should you read 0 out of 31 on risky actions? The error rate when the routes agree is 5/185 ≈ 2.7%, which predicts 0.84 errors in 31 tries (0.84 again, pure coincidence), so less than one. Seeing zero is entirely expected, and it says nothing about whether the gate beats 2.7%. For a rough sense of what a zero tells you, there's a standard shortcut:

after n successes with zero failures, 95% upper bound on the true error rate ≈ 3 / n

3/31 ≈ 9.7%. All I can honestly claim is that, on my data, this gate's error rate on risky actions is probably not above about one in ten. That's better than a single route, but nowhere near "safe to click."

To make "both agree" into a real gate, the two routes have to differ at the layer that actually fails. For example, have a model from a different family do one of the picks, or have the parse route caption icons in words, so both checks don't rest on the same 35B's eyes.

5. Using this on other people's GUI-agent numbers

The faceplants above apply just as well to other people's numbers. When I see "model X scores Y% on GUI tasks," I now ask three questions first:

  1. What was in the candidate list? The same model went 23/96 on icons when every candidate was icon-class-0 and 88/96 once it saw the image. If two methods were fed different amounts of information, the score is comparing the inputs more than the models.
  2. How many discordant questions sit behind the gap? For a gap like 190 vs. 188, count the discordant questions first. With fewer than 6, even if they all go one way, don't declare a winner yet.
  3. Who decided it "passed"? I hit this one too: the old driver version returned ok when typing into a Tk window, but nothing was actually typed. The new version honestly reports background_unavailable. "The tool says it worked" and "the app received it" are separate claims, so score against the target app's own log.

6. What's wrong with this experiment

  • The scenes are synthetic. They're Tk canvases with icons from two Linux icon themes. Real apps have busier icons, denser text, popups, and animations, so treat these numbers as relative comparisons only.
  • There's only one picking model. In chapter 4 I blame the correlated errors on "the same 35B." That's an inference. I didn't run a second model as a control.
  • The danger gate was never validated on its own. It's a text-only yes/no on the task description. The 36 risky actions are the ones it flagged, and I made no separate human labels to see what it missed.
  • The sample is small. 25 new scenes, 200 clicks, 36 risky actions. That's why the p-values in chapter 3 and the bound in chapter 4 are so wide.
  • The OCR only reads English. The recognizer is an English model. When I pointed it at screenshots from a handheld running a Chinese-language UI, it couldn't read most of the key labels. On non-English interfaces parse-then-pick will do worse than shown here. I didn't measure by how much.
  • Timing covers the decision only. Capture takes about 21 ms. Post-click waits and animations aren't included.

7. Next time you give an agent screenshot-based clicking, do it in this order

  1. Check whether the parser gives icons real names. If they all share one class label, don't send text-only candidates. Send the marked-up screenshot too.
  2. Confirm the model's coordinate convention. Qwen-family models answer in 0–1000 relative coordinates, and reading those as pixels misses every time. Check 5 questions first.
  3. If you need speed, use a hybrid router. Estimate mean latency with fast share × fast time + fallback share × fallback time, and to go faster, raise the fast share.
  4. When comparing two setups, count only discordant questions. Estimate with 2 × 0.5^k, and if k is under 6, don't declare a winner.
  5. Before relying on "act only when both routes agree," compute n × error rate A × error rate B over the full set and compare it with the observed joint errors. If the observed number is far higher, the routes share a weakness, so swap the model or the input on one of them.
  6. When a risky-action gate shows zero errors, report the 3 / n upper bound instead of 0%.
  7. Judge right or wrong by the target program's own log, not by the tool's return value.