Anthropic just released Claude Sonnet 5.5. It is the second model in the Claude 5.5 family, following Claude Opus 5.5. Anthropic positions it as a faster, lower-cost complement to Opus 5.5. It targets well-scoped everyday tasks, bug fixing, and polished documents, slides, and spreadsheets.
Is it deployable? Yes, It is live on the Claude Platform as claude-sonnet-5-5, plus AWS, Google Cloud, and Microsoft Azure. It is a closed-weights model, so self-hosting is not an option.
What Changed Versus Sonnet 5
Anthropic reports 4 main upgrades over Sonnet 5:
Speed: output generation is 30%+ faster, making it the fastest Sonnet to date.
Cost per task: up to 30% lower, because it needs fewer tokens and tool calls.
Writing: clearer prose, with early testers calling it a better collaboration partner.
Vision and long-horizon work: it is the first Sonnet to beat Pokémon Red using only screenshots.
Specs from the models overview: 1M-token context, 128K max output, and a June 2026 reliable knowledge cutoff. Adaptive thinking is on by default. Effort runs across 5 levels: low, medium, high, xhigh, and max.
Benchmarks
All scores below are vendor-reported in the launch post. Methodology lives in the Sonnet 5.5 System Card.
Terminal-Bench 4.0: 70.6%, versus 10.3% for Sonnet 5 and 66.4% for Opus 5.5 at Xhigh.
CursorBench 4.0: 55.5%, about 2 points below Opus 5.5 (57.8%).
FrontierCode 1.1: 52.1% at Xhigh and 46.2% at Max. GPT-6 Sol scored 49.3%.
GDPval-AA v2.1: 1844, versus 1846 for Opus 5.5 and 1449 for Sonnet 5.
OSWorld 2.1: 80.1% on computer use, close to Opus 5.5 (81.8%).
Humanity’s Last Exam: 64.5% with tools, up from 54.9%.
The Max result is lower than Xhigh for a reason. At Max, the model more often ran multi-agent code review. That sometimes caused timeouts or out-of-scope edits, which FrontierCode penalizes. Anthropic also states that Opus 5.5 remains clearly stronger on complex, open-ended work.
Pricing and Efficiency
List pricing is unchanged from Sonnet 5: $2 per million input tokens and $10 per million output tokens. Cache reads cost $0.20 and cache writes $2.50 per million. That is half of Opus 5.5 ($4/$20). Savings come from token efficiency, not a price cut.
Customer data supports this. Balyasny Asset Management measured about 121K tokens per answer versus 497K on Sonnet 5. Base44 reported 3.6 iterations per app build, where Opus 5 took 7.7. Zendesk processed tickets 20% faster.
Effort defaults differ by surface. Claude Code and the Claude apps default to Medium. The Claude Platform defaults to High.
How It Compares
FeatureClaude Sonnet 5.5GPT-6 SolGemini 3.1 Pro PreviewClaude Opus 5.5DeveloperAnthropicOpenAIGoogleAnthropicPrice per 1M tokens (input/output)$2 / $10$2 / $10 (prompts up to 272K)$2 / $12 (prompts up to 200K)$4 / $20Context window1M tokens1,050,000 tokens1,048,576 tokens1M tokensMax output128K tokens128K tokens65,536 tokens128K tokensKnowledge cutoffJune 2026April 20, 2026Not listedJune 2026Reasoning controlAdaptive thinking, 5 effort levels6 effort levels (none to max)Thinking supportedAdaptive thinking (always on)InputsText, imageText, imageText, image, video, audio, PDFText, imageFrontierCode 1.152.1% (Xhigh)49.3%Not reported54.4%GDPval-AA v2.118441487Not reported1846Chartography (no tools)61.6%53.6%Not reported64.4%Release stageGenerally availableGenerally availablePreviewGenerally availableOpen weightsNoNoNoNo
Benchmark scores are vendor-reported by Anthropic; GDPval-AA runs by Artificial Analysis. Prices are standard API list rates, verified September 28, 2026.
Interactive Explainer
#mtp-s55{background:#141413!important;color:#FAF9F5!important;font-family:ui-sans-serif,system-ui,-apple-system,”Segoe UI”,Roboto,Arial,sans-serif;border:1px solid #2b2a27!important;border-radius:14px;padding:22px;max-width:860px;margin:0 auto;box-sizing:border-box} #mtp-s55 *{box-sizing:border-box} #mtp-s55 .hd{display:flex;align-items:center;gap:10px;margin-bottom:4px} #mtp-s55 .dot{width:12px;height:12px;border-radius:50%;background:#D97757;animation:pulse 2s infinite} @keyframes pulse{0%{box-shadow:0 0 0 0 rgba(217,119,87,.6)}70%{box-shadow:0 0 0 10px rgba(217,119,87,0)}100%{box-shadow:0 0 0 0 rgba(217,119,87,0)}} #mtp-s55 h3{margin:0;font-size:20px;font-family:Georgia,”Times New Roman”,serif;font-weight:500;color:#FAF9F5!important} #mtp-s55 .sub{color:#B0AEA5;font-size:13px;margin:0 0 16px} #mtp-s55 .tabs{display:flex;flex-wrap:wrap;gap:6px;margin-bottom:16px} #mtp-s55 .tab{background:#1f1e1b!important;color:#B0AEA5!important;border:1px solid #2f2e2a!important;border-radius:999px;padding:8px 14px;font-size:13px;cursor:pointer;transition:all .2s} #mtp-s55 .tab:hover{color:#FAF9F5!important;border-color:#D97757!important} #mtp-s55 .tab.on{background:#D97757!important;color:#141413!important;border-color:#D97757!important;font-weight:600} #mtp-s55 .pane{display:none;animation:fade .35s ease} #mtp-s55 .pane.on{display:block} @keyframes fade{from{opacity:0;transform:translateY(6px)}to{opacity:1;transform:none}} #mtp-s55 .row{display:grid;grid-template-columns:110px 1fr 64px;align-items:center;gap:10px;margin:9px 0;font-size:13px} #mtp-s55 .trk{background:#26251f!important;height:14px;border-radius:7px;overflow:hidden} #mtp-s55 .bar{height:100%;width:0;border-radius:7px;transition:width .9s cubic-bezier(.2,.8,.2,1)} #mtp-s55 .val{text-align:right;font-variant-numeric:tabular-nums;color:#FAF9F5} #mtp-s55 .chips{display:flex;flex-wrap:wrap;gap:6px;margin-bottom:10px} #mtp-s55 .chip{background:transparent!important;color:#B0AEA5!important;border:1px solid #3a3934!important;border-radius:8px;padding:6px 10px;font-size:12px;cursor:pointer} #mtp-s55 .chip.on{border-color:#6A9BCC!important;color:#FAF9F5!important;background:#1b2530!important} #mtp-s55 .note{font-size:12px;color:#8f8d86;margin-top:10px;line-height:1.5} #mtp-s55 .card{background:#1c1b18!important;border:1px solid #2b2a27!important;border-radius:10px;padding:14px;margin-top:10px} #mtp-s55 label{font-size:13px;color:#B0AEA5;display:block;margin:10px 0 4px} #mtp-s55 input[type=range]{width:100%;accent-color:#D97757} #mtp-s55 .big{font-size:26px;font-weight:600;color:#D97757;font-variant-numeric:tabular-nums} #mtp-s55 .grid2{display:grid;grid-template-columns:1fr 1fr;gap:10px} #mtp-s55 .k{font-size:12px;color:#B0AEA5} #mtp-s55 .dial{display:flex;gap:6px;margin:8px 0} #mtp-s55 .lv{flex:1;text-align:center;padding:10px 4px;border-radius:8px;background:#1f1e1b!important;border:1px solid #2f2e2a!important;font-size:12px;cursor:pointer;color:#B0AEA5!important;transition:all .25s} #mtp-s55 .lv.on{background:#D97757!important;color:#141413!important;font-weight:600;transform:translateY(-2px)} #mtp-s55 .meter{height:10px;border-radius:5px;background:#26251f!important;overflow:hidden;margin:4px 0 10px} #mtp-s55 .mfill{height:100%;transition:width .6s ease} #mtp-s55 pre,#mtp-s55 code{background:#0e0e0d!important;color:#E8E6DC!important;border:1px solid #2b2a27!important;border-radius:8px;font-family:ui-monospace,Menlo,Consolas,monospace!important;font-size:12px!important} #mtp-s55 pre{padding:10px;margin:6px 0;white-space:pre-wrap;word-break:break-word} #mtp-s55 .bad{color:#E07A6B!important} #mtp-s55 .ok{color:#9DB57A!important} #mtp-s55 .ft{margin-top:18px;padding-top:10px;border-top:1px solid #2b2a27!important;font-size:12px;color:#8f8d86;display:flex;justify-content:space-between;flex-wrap:wrap;gap:6px} #mtp-s55 .ft b{color:#D97757} #mtp-s55 hr,#mtp-s55 p:empty,#mtp-s55 del,#mtp-s55 s{display:none!important} @media (max-width:640px){#mtp-s55{padding:14px}#mtp-s55 .row{grid-template-columns:84px 1fr 54px;font-size:12px}#mtp-s55 .grid2{grid-template-columns:1fr}#mtp-s55 .lv{font-size:11px;padding:8px 2px}}
Claude Sonnet 5.5, explained interactivelyTap through benchmarks, effort levels, task cost and the API migration checker.BenchmarksEffort dialCost per taskMigration checkerSelect an effort level. Sonnet 5.5 exposes 5.Reasoning depthSpeed and token savingsMeters are illustrative of the direction Anthropic describes (higher effort reasons longer and checks work more). Recommendations come from the Sonnet 5.5 migration guide.Illustrative estimate at list price ($2 input, $10 output per 1M tokens, identical for Sonnet 5 and 5.5).Input tokens per task: Output tokens per task on Sonnet 5: Token reduction on Sonnet 5.5: % (Anthropic: up to 30% lower cost per task)Sonnet 5, per taskSonnet 5.5, per taskAt 10,000 tasks per month you saveSavings come from fewer tokens and tool calls, not a lower sticker price. Your real ratio depends on workload.Pick a setting from your Sonnet 5 era code to see what Sonnet 5.5 returns.Sources: Anthropic launch post and Claude Platform docs, Sept 28, 2026© Marktechpost
(function(){ var R=document.getElementById(‘mtp-s55′);if(!R)return; function q(s){return R.querySelector(s)} function el(t,c,x){var e=document.createElement(t);if(c)e.className=c;if(x!==undefined)e.textContent=x;return e} function rs(){try{if(window.parent!==window)parent.postMessage({mtpH:document.body.offsetHeight+40},’*’)}catch(e){}} Array.prototype.forEach.call(R.querySelectorAll(‘.tab’),function(t){t.addEventListener(‘click’,function(){Array.prototype.forEach.call(R.querySelectorAll(‘.tab’),function(x){x.classList.remove(‘on’)});Array.prototype.forEach.call(R.querySelectorAll(‘.pane’),function(x){x.classList.remove(‘on’)});t.classList.add(‘on’);q(‘#p-‘+t.getAttribute(‘data-p’)).classList.add(‘on’);setTimeout(rs,50)})}); var C={s55:’#D97757′,s5:’#6b6a64′,o55:’#E8B59F’,g:’#6A9BCC’}; var B=[ {n:’Terminal-Bench 4.0′,u:’%’,max:100,d:[[‘Sonnet 5.5′,70.6,’s55’],[‘Sonnet 5′,10.3,’s5’],[‘Opus 5.5′,66.4,’o55′]],t:’Agentic coding in a command line. Opus 5.5 reported at Xhigh effort. GPT-6 Sol not reported.’}, {n:’CursorBench 4.0′,u:’%’,max:100,d:[[‘Sonnet 5.5′,55.5,’s55’],[‘Sonnet 5′,34.1,’s5’],[‘Opus 5.5′,57.8,’o55′]],t:’Ambiguous multi-file tasks from real Cursor sessions. GPT-6 Sol not reported.’}, {n:’FrontierCode 1.1′,u:’%’,max:100,d:[[‘Sonnet 5.5′,52.1,’s55’],[‘Sonnet 5′,42.4,’s5’],[‘Opus 5.5′,54.4,’o55’],[‘GPT-6 Sol’,49.3,’g’]],t:’Mergeable code changes. Sonnet 5.5 shown at Xhigh (52.1%). At Max it scored 46.2%.’}, {n:’GDPval-AA v2.1′,u:”,max:2000,d:[[‘Sonnet 5.5′,1844,’s55’],[‘Sonnet 5′,1449,’s5’],[‘Opus 5.5′,1846,’o55’],[‘GPT-6 Sol’,1487,’g’]],t:’Elo-style score on real work across 44 occupations, run by Artificial Analysis.’}, {n:’OSWorld 2.1′,u:’%’,max:100,d:[[‘Sonnet 5.5′,80.1,’s55’],[‘Sonnet 5′,57.0,’s5’],[‘Opus 5.5′,81.8,’o55′]],t:’Computer use, partial credit. GPT-6 Sol not reported.’}, {n:’Chartography’,u:’%’,max:100,d:[[‘Sonnet 5.5′,61.6,’s55’],[‘Sonnet 5′,15.6,’s5’],[‘Opus 5.5′,64.4,’o55’],[‘GPT-6 Sol’,53.6,’g’]],t:’Visual chart recognition, no tools.’} ]; function chipRow(box,items,fn){items.forEach(function(lbl,i){var c=el(‘button’,’chip’+(i?”:’ on’),lbl);c.type=’button’;c.addEventListener(‘click’,function(){Array.prototype.forEach.call(box.querySelectorAll(‘.chip’),function(x){x.classList.remove(‘on’)});c.classList.add(‘on’);fn(i)});box.appendChild(c)})} function drawB(i){var b=B[i],box=q(‘#bars’);box.textContent=”;var fills=[];b.d.forEach(function(r){var row=el(‘div’,’row’);row.appendChild(el(‘span’,”,r[0]));var trk=el(‘div’,’trk’);var bar=el(‘div’,’bar’);bar.style.background=C[r[2]];bar.setAttribute(‘data-w’,r[1]/b.max*100);trk.appendChild(bar);fills.push(bar);row.appendChild(trk);row.appendChild(el(‘span’,’val’,r[1]+b.u));box.appendChild(row)});q(‘#bnote’).textContent=b.t+’ Vendor-reported by Anthropic.’;setTimeout(function(){fills.forEach(function(x){x.style.width=x.getAttribute(‘data-w’)+’%’})},30);rs()} chipRow(q(‘#bchips’),B.map(function(b){return b.n}),drawB);drawB(0); var E=[ [‘low’,20,95,’Fastest and cheapest. Suggested for chat and latency-sensitive work (with medium). On several benchmarks, Low or Medium beats the best Sonnet 5 score for about a tenth of the cost per task.’], [‘medium’,40,78,’Default in Claude Code and the Claude apps. Suggested start for well-specified agentic coding and multistep tool use.’], [‘high’,60,58,’Default on the Claude Platform API. Suggested start for most workloads and harder agentic tasks. The lowest thinking mode, between_tools, works up to here.’], [‘xhigh’,80,38,’Longer reasoning. Sonnet 5.5 posted its best FrontierCode score (52.1%) here. between_tools returns a 400 error at this level.’], [‘max’,100,20,’Deepest reasoning and highest cost per task. On FrontierCode it scored lower (46.2%) because it sometimes made out-of-scope edits.’] ]; var dl=q(‘#dial’);var lvs=[]; E.forEach(function(e,i){var d=el(‘div’,’lv’,e[0]);d.addEventListener(‘click’,function(){setE(i)});dl.appendChild(d);lvs.push(d)}); function setE(i){lvs.forEach(function(x,j){x.classList.toggle(‘on’,i===j)});q(‘#m1′).style.width=E[i][1]+’%’;q(‘#m2′).style.width=E[i][2]+’%’;var box=q(‘#einfo’);box.textContent=”;var h=el(‘b’,”,’effort: “‘+E[i][0]+'”‘);h.style.color=’#D97757’;box.appendChild(h);var p=el(‘div’,”,E[i][3]);p.style.fontSize=’13px’;p.style.lineHeight=’1.6′;p.style.marginTop=’4px’;box.appendChild(p);rs()} setE(2); function cost(){var i=+q(‘#rin’).value,o=+q(‘#rout’).value,r=+q(‘#rred’).value/100;q(‘#lin’).textContent=Math.round(i/1000)+’K’;q(‘#lout’).textContent=Math.round(o/1000)+’K’;q(‘#lred’).textContent=Math.round(r*100);var a=i/1e6*2+o/1e6*10,b=a*(1-r);q(‘#c5’).textContent=’$’+a.toFixed(3);q(‘#c55’).textContent=’$’+b.toFixed(3);q(‘#csave’).textContent=’$’+Math.round((a-b)*10000).toLocaleString(‘en-US’)} [‘#rin’,’#rout’,’#rred’].forEach(function(s){q(s).addEventListener(‘input’,cost)});cost(); var M=[ [‘thinking: disabled’,’thinking={“type”: “disabled”}’,0,’400 invalid_request_error’,’thinking={“type”: “between_tools”} # effort low/medium/high only’], [‘forced tool_choice’,’tool_choice={“type”: “tool”, “name”: “get_weather”}’,0,’400: tool_choice “tool” and “any” not supported’,’tool_choice={“type”: “auto”} # mark tools strict: true’], [‘budget_tokens’,’thinking={“type”: “enabled”, “budget_tokens”: 10000}’,0,’400: thinking.type.enabled not supported’,’thinking={“type”: “adaptive”}, output_config={“effort”: “high”}’], [‘temperature / top_p’,’temperature=0.2′,0,’400 on non-default sampling values’,’# remove temperature, top_p, top_k’], [‘no thinking field’,’messages=[…] # no thinking key’,1,’200 OK, adaptive thinking runs by default’,’# read content blocks by type, pass thinking blocks back unchanged’], [‘model ID’,’model=”claude-sonnet-5″‘,1,’Swap the ID, same price’,’model=”claude-sonnet-5-5″‘] ]; function setM(i){var m=M[i],box=q(‘#mout’);box.textContent=”;box.appendChild(el(‘div’,’k’,’You send’));box.appendChild(el(‘pre’,”,m[1]));box.appendChild(el(‘div’,’k’,’Sonnet 5.5 returns’));var r=el(‘div’,m[2]?’ok’:’bad’,(m[2]?’u2713 ‘:’u2717 ‘)+m[3]);r.style.fontSize=’14px’;r.style.fontWeight=’600′;r.style.margin=’6px 0′;box.appendChild(r);box.appendChild(el(‘div’,’k’,’Use instead’));box.appendChild(el(‘pre’,”,m[4]));rs()} chipRow(q(‘#mchips’),M.map(function(m){return m[0]}),setM);setM(0); window.addEventListener(‘load’,rs);window.addEventListener(‘resize’,rs);setTimeout(rs,300); })();
Key Takeaways
Sonnet 5.5 scores 70.6% on Terminal-Bench 4.0, up from 10.3%.
Pricing stays $2/$10, but cost per task drops up to 30%.
It lands within 2 points of Opus 5.5 on GDPval-AA.
First Sonnet with cyber safeguards and reasoning-extraction classifiers.
Deployable now via API, AWS, Google Cloud, and Azure.
Check out the official announcement, migration guide, and system card. All credit goes to the researcher of this project. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.
Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us
The post Anthropic Releases Claude Sonnet 5.5: 70.6% on Terminal-Bench 4.0 at the Same $2/$10 Price appeared first on MarkTechPost.
