Test results

Claude Haiku 5.5 vs GPT-6 Luna.

Same 256 jobs in Ask Buoy. They tied on getting it right. Haiku was faster. Luna cost less.

Updated October 9, 2026 · by the Buoy team

Tie

on jobs done right

1.3 s

sooner first answer with Haiku

29%

less per job with Luna

The results

Same app. Same 256 jobs. Two models.

Claude Haiku 5.5 came out on October 7. Ask Buoy runs on GPT-6 Luna today. So we put them side by side, both at default settings.

Jobs done right

higher is better

Claude Haiku 5.593.8%
GPT-6 Luna94.1%

First answer

lower is better

Claude Haiku 5.51.6 s
GPT-6 Luna2.9 s

Whole job

middle case, lower is better

Claude Haiku 5.520.9 s
GPT-6 Luna27.9 s

Cost per job

lower is better

Claude Haiku 5.50.62¢
GPT-6 Luna0.44¢

Each job ran once. So a gap of 1 or 2 points can be luck.

By area

Each one has its strengths.

Haiku did best on long bug stories with many steps. That was the biggest area, with 56 jobs. Luna did best on web calls, app data, location and undo.

3

areas Haiku won

  • Bug stories and long chats · 56 jobs94 vs 88
  • Users and time · 16 jobs100 vs 94
  • Errors and logs · 8 jobs100 vs 88

7

areas Luna won

  • Web calls · 28 jobs96 vs 100
  • Screens and links · 22 jobs93 vs 95
  • App data stores · 21 jobs90 vs 95
  • Saved data · 18 jobs94 vs 100
  • Setup and content · 14 jobs79 vs 82
  • Location and permissions · 13 jobs85 vs 96
  • Undo and speed · 9 jobs89 vs 100

4

areas tied

  • Server data · 19 jobs100%
  • Restore points and app life · 16 jobs94%
  • Tapping and typing · 8 jobs100%
  • Basics · 8 jobs100%

On 140 jobs we kept apart for a final check, Haiku got 95.0% and Luna got 94.7%.

Turn it up

What about more thinking?

We also ran Haiku with its thinking effort set to high. It got about 1 point more right. But it was slower, and it cost more.

95.1%

jobs done right

was 93.8%

24.8 s

a whole job

was 20.9 s

0.73¢

a job

was 0.62¢

We kept the main test at default settings, so the match stays fair.

The bill

All 512 runs cost $2.72.

Both models stayed under 1¢ a job. Luna cost about 29% less.

$1.59

Claude Haiku 5.5, all 256 jobs

$1.13

GPT-6 Luna, all 256 jobs

So which one?

Pick what you care about most.

Want speed?

Claude Haiku 5.5

First answer in 1.6 s. A whole job in 20.9 s.

Want to spend less?

GPT-6 Luna

0.44¢ a job. About 29% less than Haiku.

Ask Buoy is the AI chat inside Buoy. It works in your running app. It reads the screen, checks web calls, changes app data and taps buttons.

Every number

The fine print.

We make Buoy. So here is how we tested and every score.

How we tested

  • 256 jobs in our own food ordering test app, on iOS Simulators.
  • Each job types what a person would type. Then a script checks the app itself.
  • A job counts as right only when the app shows the result.
  • Both models used their default settings. Neither got extra thinking time.
  • Both ran the same test code, the same jobs and the same checks.
  • Over two days we ran more than 2,500 tests. The final head-to-head was 512 runs.

Every area

Each area shows the share of jobs done right. The better score is in green.

AreaJobsClaude Haiku 5.5GPT-6 Luna
Bug stories and long chats5694%88%
Web calls2896%100%
Screens and links2293%95%
App data stores2190%95%
Server data19100%100%
Saved data1894%100%
Restore points and app life1694%94%
Users and time16100%94%
Setup and content1479%82%
Location and permissions1385%96%
Undo and speed989%100%
Errors and logs8100%88%
Tapping and typing8100%100%
Basics8100%100%

Fair notes

  • We make Buoy, and we wrote the jobs. Your app may give different results.
  • Each job ran once in the final test, so a gap of 1 or 2 points can be luck.
  • Luna and Haiku ran at different times, on the same test code and the same app.
  • Costs come from the tokens each run used, at list prices.

Which model does Ask Buoy use?

The hosted Ask Buoy beta runs on GPT-6 Luna today. Ask Buoy can run on either model.

So which one is better?

At default settings they got the same share of jobs right. Haiku was faster. Luna cost less. Pick the one that fits what you care about most.

Can I see the raw runs?

We kept the log of every run. Contact us if you want to check a number.

sources

Capability claims on this page come from each vendor's own documentation, read on the date shown. We did not install and run every tool listed.