Your workload should choose the inference stack.
Give InferCrane the workload, quality bar, SLOs, privacy boundary, and budget. We search models, runtimes, configurations, kernels, hardware, and supply. Then we prove the strongest qualified path behind one stable API.
- OpenAI
- Claude
- Qwen
- DeepSeek
- Kimi
- GLM
- vLLM
- NVIDIA
Scroll the board sideways →
Illustrative candidates show the decision flow. Reduced motion is on. Showing the selected example.
The model is only the start. The workload defines the system.
A model name does not tell you how it should run. Prompt shape, tool use, latency, quality, privacy, and traffic change which bottleneck matters and which execution path earns its place.
Runtimes and accelerator families trade compute, memory, software maturity, availability, and cost differently. Coupling an application permanently to one stack turns yesterday's decision into tomorrow's constraint.
The final shape is an execution layer, not another GPU picker: one stable application contract above a system that can measure, optimize, requalify, route, and recover as models, traffic, and hardware evolve.
InferCrane treats execution as a search problem.
The dominant resource changes with the workload. The serving system should be allowed to change with it without changing the application contract.
Eight layers, each one a decision.
Workload × model × runtime × configuration × decoding × kernel × silicon × cloud. Every layer is a decision InferCrane makes for your workload. Click through the stack.
01Workload
Layer 01 / 08The search starts from traffic shape, not from a vendor list. Prompt lengths, output lengths, burstiness, and latency sensitivity each point to different regions of the execution space.
We don't guess. We prove.
Six stages, every engagement, every time. This list is the legend for the execution board at the top of the page. The signal line traces how far through the method a candidate has traveled, and nothing reaches deploy without the full descent.
- 01
Profile
Capture the workload shape and a baseline: prompt and output lengths, burstiness, latency sensitivity, current cost per request.
- 02
Search
Enumerate plausible candidates across model, runtime, configuration, decoding, kernels, and silicon. Keep only what the profile makes promising.
- 03
Build
Materialize reproducible configurations, pinned and versioned. Optional kernel candidates are built only where the profile justifies the work.
- 04
Measure
Run workload-realistic tests for latency, throughput, cost, and quality with the measurement plan agreed before anything runs.
- 05
Prove
Reject failures and reject insufficient evidence alike. A configuration becomes qualified only with results that reproduce and are stored.
- 06
Deploy
The approved recipe goes behind a stable endpoint, versioned and restorable. Deployment is an approval, never an experiment side-effect.
Below the API, the entire execution stack can change.
An open-weights model does not dictate how it is served, decoded, or which silicon it runs on. In a scoped dedicated engagement, each controllable layer becomes a candidate to choose, tune, and prove. Supplier-backed APIs retain their provider-owned stack.
- Model weightsOpen models · customer checkpointsThe same weights can run on many paths
- Serving engineOur runtime profilesSwappable per recipe
- Scheduling + cacheBatching · prefix reuse · admissionWhere throughput is won or lost
- DecodingStandard · speculative · MTPLatency structure of decode
- KernelsOur kernel layer · hardware-nativeHow much silicon is actually used
- SiliconCandidate targets · qualification requiredDifferent strengths per workload
- Workload requirements
- Candidate planning + experiments
- Serving engines + scheduling
- Decoding strategies
- Kernel layer
- Recipe versioning + releases
Scope depends on the execution boundary. InferCrane can own requirements, planning, evidence, decisions, and releases while supplier-managed kernels remain explicitly provider-owned.
The best hardware depends on the workload.
Illustrative candidate lanes show how hardware remains a parameter. Actual support, availability, and selection require compatible, internally reproduced evidence for the exact workload tuple.
Candidate target. Public proof remains scoped to exact qualified tuples.
Candidate target. Workload-matched qualification is required before selection.
Candidate target. Workload-matched qualification is required before selection.
One request runs entirely on one proven target. Simple to reason about, simple to prove.
Cross-vendor prefill/decode placement is a research direction, not a currently qualified production capability.
There is no universal winner.
This illustrative comparison shows the decision InferCrane produces: candidates must clear the workload requirements before latency, throughput, and cost determine the useful frontier.
| Candidate | p50 | €/1k |
|---|---|---|
C-01 · Qwen3.5-27B-FP8 · SGLang NEXTN · 3 steps · top-k 1 · 4 draft · 1× H100 | 240 ms | €1.90 |
C-02 · GLM-5-FP8 · SGLang 947927bdb EAGLE · DSA sgl-kernel · TP8 · 8× H200 | 420 ms | €0.90 |
C-03 · DeepSeek-V4-Flash · SGLang DSpark · FlashInfer MXFP4 · TP4 · 4× H200 | 380 ms | €0.70 |
C-04 · Kimi-K2.6 · SGLang 0.5.9 MLA · tool/reasoning parsers · TP8 · 8× H200 | 610 ms | €1.40 |
Promotion follows the selected objective across the qualified frontier. Your workload decides the winner. Real campaigns use measurements from your exact workload and deployment target.
Compare routes on the same workload.
Set your traffic once. Then compare the price, cache behavior, retries, and fixed cost of two routes.
One workload
Applied to both routesCurrent route
Enter your ratesProposed route
Illustrative inputsThese are editable planning assumptions, not measured savings. InferCrane validates the real workload before recommending a route.
Open models today. Universal inference tomorrow.
Start with one API. Expand into a system that continuously improves execution across models, runtimes, kernels, clouds, and chips.
- 01
Now · founding access
Open models without infrastructure complexity
Plan and deploy open models in infrastructure you control, connect an existing endpoint, or evaluate an open alternative to the OpenAI or Claude workload you already trust.
What it unlocks
- One compatible API
- Open model evaluation
- Managed capacity only when qualified
- Existing endpoint adoption
- Quality, privacy, latency, and cost checks
- 02
Next · optimized deployments
Performance matched to your workload
InferCrane selects and tunes models, runtimes, decoding, kernels, and hardware around your requirements, then produces a portable deployment for managed, dedicated, or your cloud.
What it unlocks
- Candidate search
- Runtime and decoding optimization
- Kernel optimization where it matters
- Portable execution recipes
- 03
Then · continuous operations
Inference that improves with change
Keep your application stable while InferCrane monitors performance, requalifies new options, manages releases, and recovers from regressions.
What it unlocks
- Evidence-aware routing
- Autoscaling and observability
- Release and rollback controls
- Drift detection and requalification
- 04
North star · universal execution
Write once. Run everywhere.
Integrate one stable InferCrane API. Run supported open models across clouds, runtimes, and chips without rewriting your application. InferCrane finds, qualifies, and operates the best execution path as the ecosystem changes.
What it unlocks
- Portable model execution
- Cross-cloud and cross-chip scheduling
- Generated kernels
- Programmable heterogeneous inference
Start with one API.
Deploy an open model in your infrastructure, bring an existing endpoint, or start with the OpenAI or Claude workload you already trust.
You do not need to choose an open model first.
Bring representative OpenAI or Claude traffic and the requirements that matter to your application. InferCrane keeps that API as the control, tests a bounded set of open-model execution paths, and recommends a move only when the evidence passes.
- Your current API stays the pinned baseline
- Quality and tool behavior are checked on your examples
- Privacy, latency, and total cost are evaluated separately
Model API catalog, guarded availability.
The catalog and usage controls are live. Managed routes remain unavailable until a supplier-backed offer has current availability, pricing, funding, and reconciliation evidence.
- OpenAI-compatible endpoint
- Prepaid credits, no subscription
- Usage limits + spending caps
- Deployable and callable states stay distinct
- Callable only with current supplier evidence
- Usage and billing visibility
Your workload, a better build.
For your own infrastructure on AWS, GCP, or Kubernetes: dedicated capacity, custom weights, hardware migration, and runtime + kernel optimization where the evidence justifies it.
- Custom deployments
- Customer-owned compute
- Candidate hardware qualified per workload
- Runtime optimization
- Kernel work where justified
- Private infrastructure
Give every agent its own computer.
Run coding agents, evaluations, and long jobs away from your laptop and credentials. Keep the workspace, put compute on standby when idle, and resume through one small interface.
- Firecracker isolation on dedicated hosts
- Durable /workspace across standby and resume
- Commands, files, previews, and lifecycle evidence
- Scoped model access without long-lived keys in the sandbox
Before you move traffic.
What InferCrane changes, what it proves, and who controls deployment.