Skip to content

Blog · July 2026

9.7× cheaper than Claude Sonnet: what we found out before partnering with Kosmik Compute

  • Jakub Liška · July 2026

Before agreeing on a partnership with the Prague-based Kosmik Compute, we tested their model on 35 real client transcripts. It came out 9.7× cheaper than Claude Sonnet, more than Kosmik advertises itself.

Before agreeing on a partnership with the Prague-based Kosmik Compute, we tested their model on 35 real client transcripts. It came out 9.7× cheaper than Claude Sonnet, more than Kosmik advertises itself. The article contains the whole test, a savings calculator and above all what your company gets out of it: data in Prague, AI operation for a fraction of the price and a team that sets it up.

First the test, then a name under the offer

Transformuj.ai officially works with Kosmik Compute, a Prague-based AI infrastructure with cheaper inference. The split of roles is simple: they hold the performance, the price per token and the data in Prague, and we build solutions on it and teach people in companies to work with AI.

A man at a laptop testing a model on Kosmik Compute infrastructure

Before we said yes, we tested their model on our own production task: 35 real client transcripts that we normally work with. Not a demo from their website, but data we care about. Only after the test did we shake on it. The order is the whole point: numbers first, then a name under the offer.

What your company gets out of it

AI operation 9.7× cheaper. Measured against Claude Sonnet on our real batch, not on a marketing slide. Kosmik itself advertises 9.1× on its website. For extraction tasks dominated by input text, which is most company routine, their offer comes out even better than how they sell it themselves.

A man at a whiteboard with a diagram drawn on it

Data stays in Prague. Kosmik states that prompts are not logged or stored and that no training happens on the data. The hardware runs in certified housing in the Czech Republic and operation is compliant with GDPR. For healthcare, law, finance and public administration this is often more important than price.

We set it up and teach people. A cheap token saves nothing by itself. The saving arises only when AI takes over a specific routine: documents, emails, content, code. That is exactly our part of the work.

A table of models on Kosmik with the price per million tokens and the saving compared to Claude Sonnet

The test: 35 transcripts, zero errors

The test took place on 27 July 2026 on the Qwen 3.6 27B model, the only one we actually tested: 35 real client transcripts, 320 346 input and 16 724 output tokens, total time 2.1 minutes.

What worked: 35 out of 35 transcripts without an error and without having to repeat the query. Zero hallucinated numbers, every detail could be traced in the source text. Structured output hit all 12 required fields on the first try. In 400 operational records the model found all three extraordinary events and added up the damage correctly. Czech without errors. For two unusable transcripts the model itself wrote that they can't be analysed, instead of making up the content, exactly the behaviour we want from an extraction tool.

Savings calculator: monthly costs of Claude Sonnet versus the Qwen model on Kosmik

Scope of the test. We verified context up to 34.7 thousand tokens out of the declared 262 thousand, the tasks were extraction tasks and we measured on the shared free tier. Copywriting is a different task, the model is built for data extraction, we won't deploy it for outgoing text. Throughput: on the free tier 98.6 tokens per second, Kosmik states 3 832 for the production tier with short chats and 1 467 for longer work. We didn't measure the production tier, these are numbers from two different environments.

Everything that runs on Kosmik

Kosmik has more models in its catalogue, we have tested one so far. For the others it is a recalculation of the public price list onto our batch, at 0.92 EUR to the dollar.

The most striking number in the table isn't ours. The coding model qwen3-coder-next comes out 29.9× cheaper than Sonnet on the same batch. For startups and development teams that is an order of magnitude that changes the budget for iterations.

Integration is configuration, not a project. The API is compatible with OpenAI and Anthropic, you switch an existing application by changing two variables.

A chart of the AI usage index in Central Europe, the Czech Republic with an index of 1.84

How much it really saves

The saving is counted in tokens per month, not in percentages.

Small volumes, up to 20 million tokens a month. The saving is tens of euros to low hundreds, more of a bonus. The main argument is operation: data in Prague, nothing is logged, costs are predictable.

The volume of a small and medium company, 20 to 100 million. The saving is already visible in the operating budget. A healthcare facility with 35 million tokens saves around 1 240 EUR a year on operation and above all: for that money the AI processes roughly 3 500 messages a month, that is hundreds of hours of staff work.

A chart of the occupation categories in which AI is used in the Czech Republic

Large volumes. Legal research at 150 million: 5 320 EUR a year. Software development with Coder Next at 60 million: 5 712 EUR a year. An e-shop generating content at 400 million: 34 896 EUR a year, the only scenario where price alone decides.

All of these are model calculations from our batch and the public price list, not a price quote. But the order of magnitude holds.

A summary of AI usage data in the Czech Republic and a comparison of two types of companies

What the data on the Czech Republic show

The Anthropic Economic Index (May 2026, 121 countries) measures how and where Claude is used. With an index of 1.84 the Czech Republic is 34th out of 121 countries, so it uses it above average. But it sends less of it to work than the world, 39.6% against 43.4%. The tool is here, it just hasn't yet flowed from personal use into company operations.

Lowest compared to the world are finance and management, exactly the places where repeated routine like reporting and approval processes sits. It isn't because it couldn't be done with AI. Nobody has set it up.

Why we went into the partnership

The test was a condition, not a formality. It passed and on top of that came things we hadn't yet seen on the first call with Kosmik. The catalogue really moves, we didn't choose a partner based on a single snapshot in time.

The overlap of Transformuj.ai and Kosmik: AI adoption and infrastructure

What next

A cheaper token is not a reason to change supplier by itself. The equation includes data locality, quality on a specific task and the volume a company really processes. We calculated it on our own data and it came out well enough that we stand behind it with our name.

Calculate it on yours.

The article was first published on LinkedIn. Original article on LinkedIn

Start here

15 minutes is enough for you to know whether it makes sense.

You tell us what is bothering you. We tell you whether we can help, roughly what it costs and where I would start in your place. No presentation.

Address
Legerova 1820/39, Praha 2

The calendar is run by Cal.com (a third party) and loads only after you click.