Home / M3

AI runtime efficiency

M3

M3 cuts what you spend on AI at three layers: the tokens each request uses, the model that serves it, and the compute it runs on. Each layer works on its own, and the savings stack when you use them together.

Where the savings come from

One source of savings at each layer.

Fewer tokens

Layer 1: less to process

Compressed context and leaner requests mean fewer tokens billed for the same answer.

50–80%

Layer 2: the right model at the right price

The same model sells for 50 to 80 percent of list depending on where you buy it, and routine requests do not need top models. M3 Router handles both.

Off-peak

Layer 3: compute when it is cheap

M3 Compute prices hours outside peak demand lower, for training and batch work that can wait.

How the layers fit together

Start with any layer and add the others later.

Start

Begin where the spend is

Most clients start with tokens and routing for API spend, or with compute for training and batch jobs.

Stack

Savings compound

Each layer cuts cost at a different point: fewer tokens, sent to cheaper models, on lower-priced compute.

Dashboard

One view across M3

Usage and savings for every layer you use appear on the same dashboard.

What we measure

The same numbers across the M3 family.

MetricUnitWhy it matters
Serving cost$ per 1M tokensThe number finance asks for.
Tokens per requesttokensShows what context compression saves.
Cost per GPU-hour$ / GPU-hourReported separately for reserved and off-peak capacity.
Effective uptime%Held up by multi-pool redundancy and failover, not one provider's SLA.
Tail latencyp95 msWhat users actually feel.
Quality deltaeval scoreChange on your own test set, reported with every routing change.

Tell us what you run and what it costs today.

Start a conversation