Test results

Your “free” AI tool is costing you money.

In our test, the free tool took 2.2× longer than Buoy and ran up 2.2× the AI bill.

Updated October 9, 2026 · by the Buoy team

2.2×

faster than Argent

½

the AI bill of Argent

6×

faster than no tools

The race

Same AI. Same 14 jobs. Three toolkits.

Claude Sonnet 5.5 did 14 jobs in our test app, 3 times each. It went once with no tools, once with Argent and once with Buoy MCP.

Time

all 14 jobs, min:sec

Buoy2:56
Argent6:26
No tools17:44

AI bill

every run

Buoy$2.60
Argent$5.71
No tools$5.92+
Runs passedBuoy 41 of 42Argent 40 of 41No tools 31 of 42

No tools cost at least $5.92, because 10 runs ran out of time. Buoy used its Pro tools. We tuned Buoy on these jobs.

The film

Step right up. Admission is free.

The whole race as a two minute carnival horror film.

the race, in under two minutes

Why Buoy is faster

Buoy is already inside your app.

Other tools

Read the screen. Tap. Read the screen again.

They work from the outside, so the AI guesses from what it can see.

Buoy

Read the app's data and web calls. Right away.

So the AI makes fewer guesses and burns fewer tokens.

What's your hour worth?

Seconds add up fast.

A senior mobile dev makes about $120,000 to $160,000 a year. A lead can make up to $200,000. Here is what this test adds up to, next to Argent.

Small team

6 devs · 2 web, 2 mobile, 2 backend

$10,857

a year · 10 hours of AI waiting back a month

Mid-size team

25 devs · a few teams

$45,238

a year · 42 hours of AI waiting back a month

Big company

75 devs · 50 to 100 devs

$135,715

a year · 125 hours of AI waiting back a month

Estimates, not a promise. Each dev's AI does 20 app jobs a day, 20 work days a month. A dev costs about $72 an hour (a $150,000 salary). Each job saves the average from this test: about 15 s and 8¢ vs Argent. Next to no tools, the small team gets back about 42 hours and $3,238 a month.

Buoy Pro

Priceless for your team.

Those numbers only count your AI agent. Pro gives your whole team the same tools: in the app, on Buoy Desktop and through your AI. We can't put a price on that yet.

Support

See the app as the stuck customer, in the real production app.

QA

Jump the clock, load a saved state and retest on a real phone.

Product

Try a failed payment or a future date with no special build.

Dev

Set up a bug in one step, then hand it to the AI.

Every person who stops waiting on a dev adds up. That is the multiplier.

Free: your AI reads the screen, taps, swipes, types and reads logs. Pro: it sees inside your app.

Every number

The fine print.

We make Buoy. So here is how we tested, every job and where Buoy lost.

Every job

This is the middle of 3 runs. The fastest time for each job is in green. About 150 s means at least 2 of the 3 runs ran out of time.

JobNo toolsArgentBuoy MCP
Read one number off the screen5.5 s46.4 s7.1 s
Add 3 burgers to the cart150.3 s20.8 s15.2 s
Find why checkout gets declined112.2 s28.8 s20.6 s
Set the crowns to 1,000,00081.5 s40.4 s16.9 s
Read the last web call31.6 s16.0 s6.8 s
Save a test and run it twice150.3 s103.8 s44.8 s
Tap a tab30.7 s11.7 s5.6 s
Type a name75.7 s29.1 s9.4 s
Scroll to the end71.8 s14.7 s11.3 s
Pinch and turn the map39.2 s14.2 s7.9 s
Turn the phone sideways150.2 s6.0 s4.3 s
Record 3 taps132.8 s21.2 s12.9 s
Say what's under a finger15.0 s7.8 s6.2 s
Spot what changed on the screen16.9 s24.8 s6.8 s

Buoy was faster than Argent on all 14 jobs.

Where Buoy lost

Easy jobs don't need a tool. To read one number off the screen, no tools took a screenshot and read it in 5.5 s. Buoy took 7.1 s, and Argent took 46.4 s. No tools was also faster than Argent at spotting what changed on the screen: 16.9 s vs 24.8 s. Buoy took 6.8 s.

How we tested

  • We used Claude Sonnet 5.5 in Claude Code. It worked on our own food ordering test app on iOS Simulators. Buoy is in that app.
  • We used Argent 0.27.0 and Buoy MCP 7.0.63. Every AI also had a shell.
  • Each job ran 3 times for each tool, with a 150 second limit. A script checked the app after each run.
  • Times are the middle of the 3 runs. The totals add those up.

The AI bill

With no tools, 10 runs ran out of time. They logged $0, because they were stopped before the bill came in. But they still used tokens. So we priced their tokens from the run logs, at Sonnet prices. That came to $1.68. The same math on the runs that finished matched the real bill within 0.2%. So no tools cost at least $5.92.

Fair notes

  • We planned 15 jobs. We dropped one, because an old code change in the app had already done it for all three.
  • One Argent run was not scored, because our checker failed. So Argent has 41 runs, not 42.
  • We tuned Buoy on these jobs. Keep that in mind.
  • The race used Buoy's free tools and its Pro tools. We have not timed Free alone.
  • The team numbers are a guess built from this test. They are not measurements.
  • Argent is free and open source. Argent is a trademark of its owner.

What is free in Buoy MCP now?

Reading the screen, taps, swipes, typing, steps, screenshots, turning the phone, logs and the list of web calls. Pro tools see inside your app, like its data and web call bodies.

Can I use Buoy and Argent together?

Yes. In our tests, Argent ran in an app with Buoy installed.

Can I see the raw runs?

We kept the log of every run, with the tool calls, times and costs. Contact us if you want to check a number.

sources

Capability claims on this page come from each vendor's own documentation, read on the date shown. We did not install and run every tool listed.