Technology

The optimization layer between models and silicon.

More and more, what limits a model in production is not raw compute. It is how well its kernels fit the hardware. We build the tools that measure that fit and the system that improves it, on NVIDIA and AMD.

The problem

The limit is optimization, not compute.

A modern accelerator can deliver far more than most workloads get out of it. Three things stand in the way.

Totals hide the cause

Hardware counters report totals over a whole kernel launch. They show that a kernel is slow, not which stage is waiting on which.

Hardware keeps changing

NVIDIA, AMD and mixed fleets each need their own execution plan. Tuning done by hand for one architecture does not carry over to the next.

Optimizations interact

Changes to kernels, memory and runtime affect one another. They have to be planned together, and manual tuning loops do not scale to that.

The stack

An API is the top of the stack. We build the rest of it too.

Select a layer. The models live today are served through their providers; the layers below are what open models on our own GPUs are tuned with.

Layer 03

A planner that keeps optimizing, instead of tuning once.

It discovers optimization opportunities from how workloads really execute, models how kernels, memory and runtime interact, and plans changes that do not conflict with each other.

Architecture discovery Optimization graph Conflict-aware scheduling Continuous adaptation

Mars Optimization Brain

A system that keeps optimizing, where others tune once.

Three coordinated layers turn what the hardware and the workload are doing into changes that are safe to apply.

Architecture discovery

Maps the topology, memory hierarchy, interconnect behaviour and execution constraints of each target environment.

Optimization planning

Builds an optimization graph and generates conflict-aware plans across kernels, runtime policy and scheduling paths.

Continuous adaptation

Applies each optimization, measures its effect, and keeps improving as architectures, drivers and serving patterns change.

What this means for the API

Where the stack meets your requests.

The API and the optimization stack come from one team. This is exactly where they meet today.

Live today

Provider-served models

Qwen-SEA-LION v4.5 and GPT-6 Astra run on their providers' own infrastructure. Through MarsCompute you get one API for them, with scoped keys, spending limits and per-request metering.

Next

Models on our GPUs

Kimi K3 and GLM-5.3 will run on GPU capacity we operate. That is where this stack applies in full: we control the kernels, the runtime and the hardware they run on.

A model to call, or a fleet to tune?

Request access to the API, or talk to the optimization team about your own kernels and hardware.