All posts

Sonnet 5.5 vs Sonnet 5: Token Cost and Performance Comparison

What changes from Sonnet 5 to Sonnet 5.5? We examine side-by-side motion graphics and visual coding demos, published benchmarks, and token costs, with practical guidance on what to test before switching.

5 min read
On this page

Sonnet 5.5 promises faster results and improved output quality. Anthropic claims it delivers responses over 30% quicker while reducing task costs by up to 30%, all while maintaining identical per-token pricing.

In this blog, we run these two models side-by-side to see how both models performs across day to day tasks in a timed fashion (to measure the increased speed).

TLDR: Sonnet 5.5 is a worthy upgrade over the Sonnet 5 considering the pricing is the same, and the performance ~30% better.

Motion graphics

Yes you heard it right. These models are capable of studio-quality video generation for any prompt.

Prompt: Make a dynamic 15-second motion graphics video that shows what an incredible motion designer you are, like it's your showreel for a résumé. Go all out.

Both recordings play at their original speed. Sonnet 5's last frame is held for about 2.43 seconds so the longer Sonnet 5.5 clip can finish. This comparison shows the visual capabilities of Sonnet 5.5 far exceed the primitve visualities of Sonnet 5.

A flock of 400 starlings

The task is to create a starling murmuration in a single HTML file. A murmuration is that flowing cloud of birds that changes shape without falling apart.

Prompt: A murmuration of 400 starlings in one HTML file

Sonnet 5.5 gets its flock moving while Sonnet 5 is still writing code. Once both outputs appear, the older model's birds are more scattered in the displayed sequence. The newer model produces a denser, flowing formation, with visible trails that emphasise the movement.

That distinction matters. Drawing 400 moving objects is one part of the request. Making them look like a flock is the other. In this example, Sonnet 5.5 communicates the idea more convincingly.

Wind shaping sand dunes

Here, both models build an animated dune scene in one HTML file.

Prompt: Wind shaping sand dunes in one HTML file

Sonnet 5 produces a sparse landscape of wavy outlines. Sonnet 5.5 adds more pronounced ridges, hatching, drifting particles, and a warm sun. The extra detail gives the scene depth and makes the desert setting easier to recognise.

The newer output also appears earlier in the replay. So the trade-off on display is quite favourable: more visual detail, less waiting.

This is a design observation. A convincing desert animation does not establish that either model has built an accurate simulation of sand physics.

A clock made of clocks

The final task arranges 24 small clocks into a larger time display. Their hands need to work together to form readable digits.

Prompt: A clock made of 24 small clocks in one HTML file

This is the clearest visual difference of the four. Sonnet 5's output keeps prominent circular faces and cross-like hands, which make the larger digits harder to pick out. Sonnet 5.5 uses lighter faces and longer, connected-looking strokes. Your eye reads the time first and notices the small clocks second.

It is a better hierarchy for this particular brief. The clocks are the mechanism; the time is what the viewer needs to read.

What the benchmarks say

The broader results point in the same direction:

Sonnet 5.5 vs Sonnet 5 benchmark comparison

These are published scores from Anthropic's release table, not our measurements. GDPval-AA uses a rating rather than a percentage; these rows cannot be averaged into one score.

Faster is useful!

The price stays put

Sonnet 5.5 vs Sonnet 5 same pricing

Both models cost $2 per million input tokens and $10 per million output tokens. Anthropic attributes the lower per-task cost to fewer tokens, rather than a lower token rate. Your saving will depend on the workload.

For a repeated workflow, measure the whole job: output quality, retries, elapsed time, and cost. A quicker first response is useful only if it does not create more cleanup afterwards.

Which one should you use?

Sonnet 5.5 is the stronger candidate to try first for visual prototypes and everyday coding. These examples show an appealing combination of quicker delivery and more considered presentation.

If Sonnet 5 already powers a working process, give both models a small batch of your actual tasks before changing it. Keep the prompts and tools consistent, then inspect the failures as carefully as the successful outputs.

For agent teams, that evaluation extends to the handoffs between tasks. Oasis's Orchestrating intelligence makes the related case for shared context and human oversight. A stronger model can improve one step; the complete workflow still has to hold together.

The four tests give Sonnet 5.5 a convincing first showing. Your own repeatable work should decide whether it earns the permanent spot.

Frequently Asked Questions

How does Claude Sonnet 5.5 compare to Sonnet 5 in speed and cost?

Sonnet 5.5 generates responses over 30% faster and reduces per-task costs by up to 30% through improved token efficiency, while maintaining the exact same per-token rate as Sonnet 5.

What visual coding differences are observed between the two models?

The tests show Sonnet 5.5 creates richer, more cohesive animations, such as denser bird flocking, detailed sand dunes with lighting, and clearer visual hierarchy in complex multi-clock displays.

Should existing Sonnet 5 users immediately switch to Sonnet 5.5?

While Sonnet 5.5 excels at visual prototyping and benchmark scores, teams should test both models on a small batch of actual tasks to verify workflow reliability before making permanent switches.

Last updated: Oct 1, 2026

Build your agent team in 30 seconds.

Build agent teams that work along with your team. Free to start, no card required.