Samsa Benchmarks
Vectorization benchmark: which AI turns an illustration into an editable SVG?
The Samsa vectorization benchmark measures how well AI models and vectorization tools turn a raster illustration into an SVG a designer can edit. Every participant gets the same prompt and the same two images. Each output is scored on fidelity to the original, on structure (named, complete layers) and on focal detail. Cost and time per image are shown beside the score, never inside it, and every run's raw numbers are published.
Score against cost
Every participant, placed by its Samsa Vectorization Score and what one image cost. Up and to the left is better; the line joins the participants nobody beats on both.
- Track
- Score
- Fidelity
- Structure
- Focal detail
- Cost per image (USD)
- Time per image
- Parity gate
- Anthropic
- OpenAI
- Agentic
- Single-shot
- Fixed loop
- Tools
- Pareto frontier
Leaderboard
| Claude Opus 5.5 · max | 1 | Anthropic | Claude Code 2.1.291 | Agentic | max | 82.1 | 78.5 | 92.1 | 70.6 | pass | $51.9 (—) | 118 min (—) | 2 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Claude Sonnet 5.5 · xhigh | 2 | Anthropic | Claude Code 2.1.291 | Agentic | xhigh | 79.5 | 76.0 | 89.2 | 68.8 | pass | $32.8 ($20.5–$45.5) | 92 min (84 min–102 min) | 4 |
| Claude Opus 5.5 · high | 3 | Anthropic | Claude Code 2.1.291 | Agentic | high | 78.5 | 71.7 | 92.2 | 69.1 | pass | $22.2 ($13.3–$28.0) | 72 min (40 min–95 min) | 4 |
| Claude Opus 5.5 · xhigh | 4 | Anthropic | Claude Code 2.1.291 | Agentic | xhigh | 78.4 | 72.6 | 90.7 | 68.8 | pass | $30.9 ($20.8–$50.5) | 84 min (63 min–124 min) | 6 |
| GPT 6.1 Sol · xhigh | 5 | OpenAI | Codex CLI 0.160.0 | Agentic | xhigh | 73.9 | 64.0 | 89.0 | 73.1 | pass | $8.0 ($6.0–$15.4) | 75 min (46 min–157 min) | 6 |
| GPT 6.1 Sol · high | 6 | OpenAI | Codex CLI 0.160.0 | Agentic | high | 73.3 | 60.8 | 91.1 | 73.1 | mixed | $5.9 (—) | 56 min (—) | 2 |
| Claude Sonnet 5.5 · high | 7 | Anthropic | Claude Code 2.1.291 | Agentic | high | 72.7 | 65.2 | 88.2 | 61.3 | pass | $14.3 (—) | 53 min (—) | 2 |
| Claude Opus 5.5 · medium | 8 | Anthropic | Claude Code 2.1.291 | Agentic | medium | 72.5 | 64.2 | 88.3 | 63.1 | pass | $13.1 (—) | 62 min (—) | 2 |
| GPT 6.1 Sol · medium | 9 | OpenAI | Codex CLI 0.160.0 | Agentic | medium | 68.4 | 53.7 | 88.5 | 70.6 | mixed | $5.2 (—) | 52 min (—) | 2 |
| GPT 6.1 Sol · xhigh | 10 | OpenAI | API call, no tools (Codex CLI 0.160.0) | Single-shot | xhigh | 48.8 | 15.1 | 87.2 | 67.5 | fail | $0.34 ($0.29–$0.45) | 13 min (10 min–21 min) | 6 |
| GPT 6.1 Sol · medium | 11 | OpenAI | API call, no tools (Codex CLI 0.160.0) | Single-shot | medium | 47.3 | 15.4 | 86.9 | 68.1 | fail | $0.17 ($0.12–$0.24) | 5.4 min (3.7 min–7.5 min) | 10 |
| vtracer 0.6.15, default settings | 12 | visioncortex (open source) | — | Tool | — | 43.4 | 68.9 | 2.5 | 53.8 | mixed | $0.00 (—) | 0.1 min (—) | 2 |
| Claude Opus 5.5 · medium | 13 | Anthropic | API call, no tools (Claude Code 2.1.292) | Fixed loop | medium | 38.2 | 4.0 | 84.4 | 44.4 | fail | $1.8 ($1.5–$2.4) | 11 min (7.2 min–15 min) | 4 |
| Claude Fable 5.1 · medium | 14 | Anthropic | API call, no tools (Claude Code 2.1.292) | Single-shot | medium | 37.4 | 3.1 | 80.8 | 48.8 | fail | $0.94 ($0.67–$1.3) | 3.4 min (2.5 min–4.5 min) | 6 |
| Claude Opus 5.5 · medium | 15 | Anthropic | API call, no tools (Claude Code 2.1.292) | Single-shot | medium | 34.5 | 0.0 | 80.0 | 42.5 | fail | $0.33 ($0.16–$0.42) | 2.8 min (1.7 min–3.5 min) | 6 |
| Claude Sonnet 5.5 · medium | 16 | Anthropic | API call, no tools (Claude Code 2.1.292) | Fixed loop | medium | 32.8 | 4.5 | 74.8 | 29.4 | fail | $0.38 ($0.23–$0.59) | 3.7 min (1.8 min–6.2 min) | 4 |
| Claude Haiku 5.5 · medium | 17 | Anthropic | API call, no tools (Claude Code 2.1.294) | Single-shot | medium | 31.6 | 0.2 | 77.0 | 30.6 | fail | $0.01 ($0.00–$0.01) | 1 min (0.5 min–1.4 min) | 6 |
| Claude Sonnet 5.5 · medium | 18 | Anthropic | API call, no tools (Claude Code 2.1.292) | Single-shot | medium | 30.1 | 0.0 | 74.9 | 30.0 | fail | $0.07 ($0.04–$0.15) | 1.2 min (0.5 min–2.1 min) | 6 |
| Claude Fable 5.1 · xhigh Partial: 1 of 2 images | — | Anthropic | Claude Code 2.1.291 | Agentic | xhigh | 63.1 | 52.6 | 81.6 | 55.0 | fail | $42.9 (—) | 88 min (—) | 1 |
| Gemini 3.1 Pro · max Partial: 1 of 2 images | — | Gemini CLI 0.62.0 | Agentic | max | 46.7 | 58.0 | 35.0 | 36.3 | pass | $1.6 (—) | 11 min (—) | 1 |
What does edition 2026.10 of the vectorization benchmark show?
In Samsa's vectorization benchmark, edition 2026.10 (8 October 2026), Claude Opus 5.5 at max effort scored highest, with a Samsa Vectorization Score of 82.1 out of 100, at a median cost of $43.5 for the flat laptop image and $60.2 for the textured e-bike image. The ranking rests on 87 scored runs of 23 configurations on two illustrations. Claude Opus 5.5 at max effort is a reference arm with one run per image.
- Of the AI models, only agentic runs reached visual parity. A run is agentic when the model works in its vendor's own coding agent, renders its SVG, compares it with the original and corrects it over many rounds. 29 of 32 agentic runs passed the eight-criterion parity gate. None of the 48 single-shot and fixed-loop runs did, whichever model made them.
- Tracers copy the pixels and lose the objects. vtracer and Recraft's Vectorize reproduce the colours well, and vtracer even passes the parity gate on the flat laptop image. Both score 4 or less out of 100 on structure: on the e-bike, vtracer returned 9,328 unnamed shapes and Recraft 1,462, with every object cut into the pieces that happen to be visible.
- The cheapest strong result is not a Claude model. GPT 6.1 Sol at xhigh effort scored 73.9, at a median of $6.8 and $9.1 per image. That is 8.2 points behind the top configuration for about 15 % of its cost, which puts it on the frontier line of the chart above.
- Quiver AI draws objects, not yet the picture. Arrow 2 Telos returns a few large, grouped objects, so its structure score sits between the tracers and the agentic runs (58.1 on the e-bike). Its fidelity stays far below the parity gate on both images.
- More effort buys fidelity, at a steep price. Claude Opus 5.5 at max effort reached the highest e-bike fidelity of any run (71.8) for $60.2. The effort chart below shows where each extra dollar stops moving the score.
Claude Fable 5.1 at xhigh effort, a reference arm, ran once on the laptop image and failed the parity gate there, at $42.9; it is listed as partial, and its e-bike run comes in the update. Claude Haiku 5.5 is in this edition as a single-shot arm at medium effort: six runs, $0.04 for all of them together, none at parity, an overall score of 31.6. Its agentic arms (xhigh and high) and the hybrid, a GPT 6.1 Sol draft polished by Claude Opus 5.5 at high effort, are still running and will be added in an update. Gemini 3.1 Pro ran once, on the laptop image only, and scored 46.7: it is listed as partial, without an overall rank.
Run-to-run spread
The same agentic configuration on the same image moved by up to 4.0 points between runs on the e-bike, and its cost by up to 1.7 times. That is enough to swap neighbours in the ranking. So the configurations that matter for a product decision ran up to three times, and the leaderboard shows n and the range for every row with more than one run.
Why does a traced SVG fail the structure test?
A traced SVG fails because it describes the picture as coloured areas, not as things. Samsa built this benchmark because clients ask for editable files: they want to move the rider, recolour the jacket or take the bike out of the scene. In a trace, the far leg exists only as the slivers visible between the frame and the crank, and the sky is a dozen fragments around the clouds. Nothing has a name, so a designer cannot find anything.
A rebuilt SVG is what an illustrator would hand over: named objects in a sensible layer order, each drawn complete, including the parts other objects hide. That is what the structure score rewards, and it is why structure carries 35 % of the score next to fidelity's 50 %. A file that looks right and cannot be edited does not do the job a vector file is ordered for.
What do the results mean if you need an editable SVG?
- For an icon or a logo that only has to scale, a tracer is enough. It is free or close to it and takes seconds.
- For an illustration you want to edit, plan for an agentic run: $5.3–$60.2 of compute and 31–128 minutes per image at this level, not cents and seconds.
- Do not judge by the thumbnail. A traced file and a rebuilt file can look the same at full size; open the layers before you accept either.
- One run is a sample. If the result matters, run twice and keep the better file, or plan a human pass on the focal details: faces, hands, mechanics.
The step-by-step version, with the checks a finished file has to pass, is in the guide on how to vectorize an image with AI and keep it editable.
Cite this
Samsa, «Vectorization benchmark», edition 2026.10, 8 October 2026, samsa.ai/benchmarks/vectorization/. Prompts, judge prompts and every run's raw metrics are published with the dataset.
What more effort buys
The same model at each reasoning effort: how the score moves as the cost per image climbs.
- GPT 6.1 Sol · medium → high → xhigh
- Claude Sonnet 5.5 · high → xhigh
- Claude Opus 5.5 · medium → high → xhigh → max
Results
The original, then every render in rank order. Open a card to compare it with the original and inspect the zoom crops.
Results: E-bike ride
-
Original
2400 × 1792 px
-
#1 Claude Opus 5.5 · max
76.7 Score $60.3 128 min pass Agentic
-
#2 Claude Sonnet 5.5 · xhigh
74.8 Score $42.2 100 min pass Agentic
-
#3 Claude Opus 5.5 · high
71.7 Score $26.2 80 min pass Agentic
-
#4 Claude Opus 5.5 · xhigh
73.1 Score $34.7 85 min pass Agentic
-
#5 GPT 6.1 Sol · xhigh
72.2 Score $9.1 93 min pass Agentic
-
#6 GPT 6.1 Sol · high
72.1 Score $6.4 65 min fail Agentic
-
#7 Claude Sonnet 5.5 · high
70.2 Score $20.1 75 min pass Agentic
-
#8 Claude Opus 5.5 · medium
67.4 Score $16.4 91 min pass Agentic
-
#9 GPT 6.1 Sol · medium
72.2 Score $8.2 86 min pass Agentic
-
#10 GPT 6.1 Sol · xhigh
48.0 Score $0.38 15 min fail Single-shot
-
#11 GPT 6.1 Sol · medium
45.5 Score $0.21 6.9 min fail Single-shot
-
#12 vtracer 0.6.15, default settings
36.5 Score $0.00 0.1 min fail Tool
-
#13 Claude Opus 5.5 · medium
40.0 Score $2.1 14 min fail Fixed loop
-
#14 Claude Fable 5.1 · medium
39.4 Score $1.2 4.4 min fail Single-shot
-
#15 Claude Opus 5.5 · medium
33.1 Score $0.39 3.5 min fail Single-shot
-
#16 Claude Sonnet 5.5 · medium
34.7 Score $0.52 5.2 min fail Fixed loop
-
#17 Claude Haiku 5.5 · medium
30.5 Score $0.01 1.4 min fail Single-shot
-
#18 Claude Sonnet 5.5 · medium
30.1 Score $0.11 1.8 min fail Single-shot
Original against Claude Opus 5.5 · max
Original
Claude Opus 5.5 · max Zoom crops
Face and helmet
Drivetrain
Results: Hugging the laptop
-
Original
1376 × 768 px
-
#1 Claude Opus 5.5 · max
87.5 Score $43.5 108 min pass Agentic
-
#2 Claude Sonnet 5.5 · xhigh
84.3 Score $23.4 84 min pass Agentic
-
#3 Claude Opus 5.5 · high
85.2 Score $18.1 64 min pass Agentic
-
#4 Claude Opus 5.5 · xhigh
83.7 Score $27.1 83 min pass Agentic
-
#5 GPT 6.1 Sol · xhigh
75.7 Score $6.8 57 min pass Agentic
-
#6 GPT 6.1 Sol · high
74.4 Score $5.3 48 min pass Agentic
-
#7 Claude Sonnet 5.5 · high
75.2 Score $8.5 31 min pass Agentic
-
#8 Claude Opus 5.5 · medium
77.5 Score $9.8 33 min pass Agentic
-
#9 GPT 6.1 Sol · medium
64.6 Score $2.3 18 min fail Agentic
-
#10 GPT 6.1 Sol · xhigh
49.6 Score $0.31 11 min fail Single-shot
-
#11 GPT 6.1 Sol · medium
49.2 Score $0.13 4 min fail Single-shot
-
#12 vtracer 0.6.15, default settings
50.2 Score $0.00 0 min pass Tool
-
#13 Claude Opus 5.5 · medium
36.5 Score $1.6 7.7 min fail Fixed loop
-
#14 Claude Fable 5.1 · medium
35.4 Score $0.68 2.5 min fail Single-shot
-
#15 Claude Opus 5.5 · medium
36.0 Score $0.26 2.2 min fail Single-shot
-
#16 Claude Sonnet 5.5 · medium
31.0 Score $0.23 2.2 min fail Fixed loop
-
#17 Claude Haiku 5.5 · medium
32.6 Score $0.00 0.6 min fail Single-shot
-
#18 Claude Sonnet 5.5 · medium
30.1 Score $0.04 0.5 min fail Single-shot
-
Claude Fable 5.1 · xhigh Partial: 1 of 2 images
63.1 Score $42.9 88 min fail Agentic
-
Gemini 3.1 Pro · max Partial: 1 of 2 images
46.7 Score $1.6 11 min pass Agentic
Original against Claude Opus 5.5 · max
Original
Claude Opus 5.5 · max Zoom crops
Face
Hands
Structure breakdown
The automated structure checks per participant, models and tools together. A dash is a check this edition did not run.
How does a benchmark run work?
Every participant gets the same frozen prompt and the same images, under protocol 1.0.1. One run is one fresh session of the vendor's own agent with the prompt as its only input, and no messages in between. The run iterates until the parity gate passes and the focal-detail review is clean, or until it reaches a cap of three hours or 25 scored rounds, and then leaves its best state.
- Harness per vendor, disclosed. Claude models run in Claude Code 2.1.291 (the Haiku 5.5 arms in 2.1.294, the first version that knows the model), GPT 6.1 Sol in Codex CLI 0.160.0 on the fast tier, Gemini in Gemini CLI 0.62.0 with an API key. What is measured is model plus agent, and the leaderboard names both.
- Interrupted runs are discarded and repeated. A run stopped from outside (a usage limit, a killed process, a network error) is not scored, and its cost is reported nowhere. The repeat gets the next run number.
- Sub-agents are allowed where the agent offers them. Their tokens count, and the dataset records that the run delegated.
- One machine, one toolkit. Every agentic run executes on the same Mac, and the scoring toolkit is pinned by hash and never changed mid-edition.
What are the four tracks?
A track says how a participant produced its SVG. The chart draws each track with its own mark, and the leaderboard shows it as a column.
| Track | How the SVG is made | Who is in it |
|---|---|---|
| Agentic | The model works in its vendor's agent, writes and runs code, renders and compares its SVG, and iterates under the caps. | Claude Fable, Opus and Sonnet, GPT 6.1 Sol and Gemini 3.1 Pro; Claude Haiku and the Sol + Opus hybrid join in an update |
| Single-shot | One call with the image and the prompt, no tools. The reply is the SVG. | Claude Opus, Sonnet, Fable and Haiku, GPT 6.1 Sol |
| Fixed loop | Four rounds in which the model sees its render and the parity feedback and rewrites the SVG, no tools. The best round counts. | Claude Opus 5.5 and Sonnet 5.5 |
| Tool | A vectorizer's own output, scored like any other SVG. | Recraft Vectorize, Quiver AI Arrow 2 and Arrow 2 Telos, vtracer |
Which images are in the suite?
Two illustrations made with Samsa, one per difficulty tier. Each comes with an inventory: the list of objects a complete vectorization must contain, written by Samsa and frozen for the edition.
- Laptop, tier 1 (flat): flat fills, no texture, overlapping organic shapes, a face and hands. 1376 × 768 pixels, 32 inventory objects.
- E-bike, tier 2 (textured): film grain, stipple, gradients, 56 spokes, a face, deep occlusion behind the far leg and the crank. 2400 × 1792 pixels, 53 inventory objects.
- A tier 3 scene is reserved for a later edition.
How is a run scored?
Every run gets three sub-scores from 0 to 100, combined into the Samsa Vectorization Score. A participant's number per image is the median over its runs; the overall rank is the mean of its per-image medians. A participant measured on only one image is listed as partial, without an overall rank. Cost and time are axes beside the score, never inside it.
| Part | Weight | What it measures |
|---|---|---|
| Fidelity | 50 % | Mean colour difference, structural similarity, the worst local window and, on the e-bike, the grain match, against the original with both images composited on white. The eight-criterion parity gate is shown as a badge, not folded into the score. |
| Structure | 35 % | 60 % automated: named layers, inventory coverage, completeness of hidden geometry and path economy, with penalties for embedded rasters, clip paths that fake occlusion and baked transforms. 40 % judged: object semantics, completeness under occlusion, editability, layer order and path quality. |
| Focal detail | 15 % | Two standard zoom crops per image (face and hands on the laptop, face and drivetrain on the e-bike), judged against the original's crop for likeness and craft. |
The judged parts come from two vision models of different vendors, Claude Opus 5.5 and GPT 6.1 Sol, which see anonymized, shuffled outputs. Their prompts are published with the results.
How are cost and time counted?
- Cost is API-equivalent: the tokens a run used, by class (input, output including reasoning, cache write, cache read), times the vendor's list price on the run date. The price snapshot is stored with every run (6 October 2026). It is what the run would cost on the API, not an invoice.
- Codex runs on the fast tier record the fast-tier cost and the standard-tier equivalent. Haiku 5.5 is priced per request on its two prompt-length rate cards.
- A tool counts at its published price per conversion, and its time is one timed re-run.
- Time is wall clock from the launch of the agent to its final report.
- Rows with more than one run show the median with its range. A row with one run shows the value alone, because one run has no spread.
What changed in protocol 1.0.1?
Outputs whose canvas differs from the source are now aligned by their content, not by their frame. The scorer renders the SVG, takes the bounding box of what it drew, maps it onto the source's content box with one uniform scale and pads transparently. If the proportions of the two boxes differ by more than 10 %, the whole canvas is fitted instead. The rule applies to every participant. It was added because Quiver AI's exports carry a margin of about 2.4 % that cost them nearly all their fidelity under the old fit-and-pad rule: a framing artefact, not a vectorization fault.
What are the limitations?
- Two images. The ranking says how the participants handle a flat and a textured illustration, not photographs, logos or technical drawings.
- Model plus agent. Each model runs in its own vendor's agent, so the agent is part of what is measured. A harness-independent track may follow.
- Single runs where marked. Claude Fable 5.1 at xhigh and Claude Opus 5.5 at max are reference arms with one run per image because of their per-run cost, and the lower effort rungs ran once too. Gemini 3.1 Pro ran once, on the laptop image, under a $20 API cap. Claude Opus 5.5 at high effort has two runs per image so far, and Claude Sonnet 5.5 at xhigh two on the e-bike; their third runs are under way and will be added in an update. Claude Fable 5.1 at xhigh effort has no e-bike run yet: it was interrupted by a usage limit, discarded under the protocol, and is repeated for the update.
- The completeness check on its own rewards few large objects, which is how Arrow 2 Telos draws. The structure score as a whole still orders hand-built files above Telos and Telos above traces, but read that sub-score with this in mind.
- Judges are models. They see anonymized outputs and their prompts are published, but they are not designers. Samsa spot-checks the focal crops by hand before publication.
- Samsa builds on Anthropic models, and one of the two judges is an OpenAI model. The fidelity score and the automated half of the structure score need no judge, and every raw metric is published so the result can be checked.
How can I submit a model or a tool?
Run the frozen prompt on both images under the protocol, or, for a tool, send the SVG outputs for both images with the price and a timed run. Samsa scores them with the same toolkit and adds them in the next update. Get in touch to submit one.
Changelog
-
Edition published.
Frequently asked questions about the vectorization benchmark
What does the vectorization benchmark measure?
One task: turning a raster illustration into an SVG a designer can edit. Each run is scored on fidelity to the original (50 %), on structure, meaning named, complete and economical layers (35 %), and on focal details such as faces and hands (15 %). The weighted sum is the Samsa Vectorization Score. Cost and time per image are shown beside the score and never change it.
Which AI is best at vectorizing an illustration?
In edition 2026.10, Claude Opus 5.5 at max effort ranked first with a score of 82.1, followed by Claude Sonnet 5.5 at xhigh effort (79.5). GPT 6.1 Sol at xhigh effort scored 73.9 at a fraction of the cost. Every top configuration ran as an agent that renders and corrects its own SVG; no single prompt to any model came close.
Can one prompt to an AI model turn an image into an editable SVG?
Not at parity, in this benchmark. Single-shot calls to Claude and GPT models returned clean, well-named SVGs within minutes, for cents to about a dollar. None of the 48 single-shot and fixed-loop runs passed the parity gate: the drawings are stylised versions of the picture, not copies. Showing the model its render for four rounds did not close the gap. Parity took an agent that measures and corrects its output.
Why do tracing tools score low when their SVG looks right?
Because looking right is half of the job. A tracer splits the image into flat shapes by colour: on the e-bike, vtracer produced 9,328 unnamed shapes. Its fidelity is high, but no shape is an object and every hidden part is missing. The structure score, 35 % of the total, measures exactly what a designer needs to edit the file, so tracers lose most of it.
How are cost and time counted in the benchmark?
Cost is the tokens a run used times the vendor's list price on the run date, split by token class, with the price snapshot stored with the run. It is an API-equivalent figure, not an invoice. A tool counts at its published price per conversion. Time is wall clock from launch to final report. Both are medians over a participant's runs, shown with the range where there is more than one run.
Is the vectorization benchmark independent?
No benchmark run by a company is fully independent, and Samsa builds on Anthropic models. What makes this one checkable: one frozen prompt for everyone, automated fidelity scoring, two judges from different vendors on anonymized outputs, published judge prompts, and every raw metric published with its run. Anyone can re-run the protocol and compare.
Can I submit a model or a tool?
Yes. Run the frozen prompt on both images under the protocol, or send a tool's SVG outputs for both images with its price and a timed run. Samsa scores the submission with the same toolkit and adds it in the next update. Use the contact page to get in touch.
How often is the benchmark updated?
When a relevant model or tool ships. Rows are added, never edited: a participant measured again later gets a new edition tag, price snapshots stay with their runs, and the changelog on this page records every change with its date. Claude Haiku 5.5, released while the 2026.10 runs were under way, joined the edition that way.
Generate on-brand images and export them as SVG
Samsa's image generator learns your brand. Download your images as layered SVG files and hand them to your designers.