The execution layer for open intelligence

Your workload should choose the inference stack.

Give InferCrane the workload, quality bar, SLOs, privacy boundary, and budget. We search models, runtimes, configurations, kernels, hardware, and supply. Then we prove the strongest qualified path behind one stable API.

Open console
Use the productPlan before spend · bring a model or existing endpoint
Bring the stack you use todayTechnology references · not customer or endorsement claims
  • OpenAI
  • Claude
  • Qwen
  • DeepSeek
  • Kimi
  • GLM
  • vLLM
  • NVIDIA
01 / PROFILE
01 WORKLOAD / OPEN SEARCH
MODELRUNTIMEDECODINGKERNELSILICONTTFTmsITLmsTOK/S/ userSLO OUTPUTtok/sCOST$ / 1M outWORKLOADAgentic inference4K in / 512 outstreamingOne workload.Every execution possibility.01GLM-5-FP8zai-org/GLM-5-FP8SGLang 947927bdbcandidate runtimeEAGLE · 3 steps /…candidate configurationsgl-kernel DSAcandidate backend8× H200 · TP8candidate topology—————02Kimi-K2.6moonshotai/Kimi-K2.6SGLang 0.5.9candidate runtimeAutoregressivecandidate configurationSGLang MLAcandidate backend8× H200 · TP8candidate topology—————03Qwen3.5-27B-FP8Qwen/Qwen3.5-27B-FP8SGLang maincandidate runtimeNEXTN · 3 steps /…candidate configurationSGLang nativecandidate backend1× H100 80 GBcandidate topology—————04DeepSeek-V4-Flashdeepseek-ai/DeepSeek-V4-Flash-0731SGLang cookbookcandidate runtimeDSparkcandidate configurationFlashInfer MXFP4candidate backend4× H200 · TP4candidate topology—————05DeepSeek-V4-Flashdeepseek-ai/DeepSeek-V4-Flash-0731vLLM recipecandidate runtimeDSpark · 7 tokenscandidate configurationDeepGEMM Mega MoEcandidate backend4× GB300 · DP4/EPcandidate topology—————DEPLOYSame recipe. Same evidence.ILLUSTRATIVE SELECTED RECIPEillustrative / QWEN-03ModelQwen3.5-27B-FP8 · Qwen/Qwen3.5-27B-FP8RuntimeSGLang main · candidate runtimeDecodingNEXTN · 3 steps / 4 draft · candidate configurationKernelSGLang native · candidate backendSilicon1× H100 80 GB · candidate topologyEVIDENCETTFT—msITL—msTOK/S—/ userSLO OUTPUT—tok/sCOST—$ / 1M outsynthetic screen · no qualification claimdemo:recipe/qwen-03PROFILE → SEARCH → BUILD → MEASURE → PROVE → DEPLOYInferCraneFind, prove and run the best execution pathfor every AI workload.EXECUTION, PROVEN.

Scroll the board sideways →

00.0 / 27.4s

Illustrative candidates show the decision flow. Reduced motion is on. Showing the selected example.

Phase 1 of 6: Profile
Manifesto

The model is only the start. The workload defines the system.

A model name does not tell you how it should run. Prompt shape, tool use, latency, quality, privacy, and traffic change which bottleneck matters and which execution path earns its place.

Runtimes and accelerator families trade compute, memory, software maturity, availability, and cost differently. Coupling an application permanently to one stack turns yesterday's decision into tomorrow's constraint.

The final shape is an execution layer, not another GPU picker: one stable application contract above a system that can measure, optimize, requalify, route, and recover as models, traffic, and hardware evolve.

InferCrane treats execution as a search problem.
Search, measure, qualify. Then commit.
Where the time goes
Tool-heavy agentProfile 01Long-context analysisProfile 02Batch extractionProfile 03Realtime chatProfile 04PrefillDecodeMemoryBatchingAttention

The dominant resource changes with the workload. The serving system should be allowed to change with it without changing the application contract.

Methodology

We don't guess. We prove.

Six stages, every engagement, every time. This list is the legend for the execution board at the top of the page. The signal line traces how far through the method a candidate has traveled, and nothing reaches deploy without the full descent.

  1. 01

    Profile

    Capture the workload shape and a baseline: prompt and output lengths, burstiness, latency sensitivity, current cost per request.

  2. 02

    Search

    Enumerate plausible candidates across model, runtime, configuration, decoding, kernels, and silicon. Keep only what the profile makes promising.

  3. 03

    Build

    Materialize reproducible configurations, pinned and versioned. Optional kernel candidates are built only where the profile justifies the work.

  4. 04

    Measure

    Run workload-realistic tests for latency, throughput, cost, and quality with the measurement plan agreed before anything runs.

  5. 05

    Prove

    Reject failures and reject insufficient evidence alike. A configuration becomes qualified only with results that reproduce and are stored.

  6. 06

    Deploy

    The approved recipe goes behind a stable endpoint, versioned and restorable. Deployment is an approval, never an experiment side-effect.

Deep technology

Below the API, the entire execution stack can change.

An open-weights model does not dictate how it is served, decoded, or which silicon it runs on. In a scoped dedicated engagement, each controllable layer becomes a candidate to choose, tune, and prove. Supplier-backed APIs retain their provider-owned stack.

Execution stack · top to bottom
  1. Model weightsOpen models · customer checkpoints
    The same weights can run on many paths
  2. Serving engineOur runtime profiles
    Swappable per recipe
  3. Scheduling + cacheBatching · prefix reuse · admission
    Where throughput is won or lost
  4. DecodingStandard · speculative · MTP
    Latency structure of decode
  5. KernelsOur kernel layer · hardware-native
    How much silicon is actually used
  6. SiliconCandidate targets · qualification required
    Different strengths per workload
What the platform controls
  • Workload requirements
  • Candidate planning + experiments
  • Serving engines + scheduling
  • Decoding strategies
  • Kernel layer
  • Recipe versioning + releases

Scope depends on the execution boundary. InferCrane can own requirements, planning, evidence, decisions, and releases while supplier-managed kernels remain explicitly provider-owned.

Multi-chip story

The best hardware depends on the workload.

Illustrative candidate lanes show how hardware remains a parameter. Actual support, availability, and selection require compatible, internally reproduced evidence for the exact workload tuple.

NVIDIA
H100 class
Decode latency
Cost profile
Throughput headroom

Candidate target. Public proof remains scoped to exact qualified tuples.

Illustrative fit · fastest
AMD
MI355X class
Decode latency
Cost profile
Throughput headroom

Candidate target. Workload-matched qualification is required before selection.

Illustrative candidate
INTEL
Gaudi class
Decode latency
Cost profile
Throughput headroom

Candidate target. Workload-matched qualification is required before selection.

Illustrative candidate
Single-target execution

One request runs entirely on one proven target. Simple to reason about, simple to prove.

Heterogeneous execution

Cross-vendor prefill/decode placement is a research direction, not a currently qualified production capability.

Evidence over opinionIllustrative

There is no universal winner.

This illustrative comparison shows the decision InferCrane produces: candidates must clear the workload requirements before latency, throughput, and cost determine the useful frontier.

200 ms400 ms600 ms€0.5€1.0€1.5€2.0€2.5p50 latency →€ / 1k requests →C-01C-02C-03C-04
Candidatep50€/1k
C-01 · Qwen3.5-27B-FP8 · SGLang
NEXTN · 3 steps · top-k 1 · 4 draft · 1× H100
240 ms€1.90
C-02 · GLM-5-FP8 · SGLang 947927bdb
EAGLE · DSA sgl-kernel · TP8 · 8× H200
420 ms€0.90
C-03 · DeepSeek-V4-Flash · SGLang
DSpark · FlashInfer MXFP4 · TP4 · 4× H200
380 ms€0.70
C-04 · Kimi-K2.6 · SGLang 0.5.9
MLA · tool/reasoning parsers · TP8 · 8× H200
610 ms€1.40

Promotion follows the selected objective across the qualified frontier. Your workload decides the winner. Real campaigns use measurements from your exact workload and deployment target.

Cost comparison

Compare routes on the same workload.

Set your traffic once. Then compare the price, cache behavior, retries, and fixed cost of two routes.

One workload

Applied to both routes

Current route

Enter your rates

Proposed route

Illustrative inputs
Current monthly
€14,134.00
Proposed monthly
€8,439.12
Estimated change
−€5,694.88
40.3% lower
Cost per request
€0.00707 → €0.00422
current → proposed

These are editable planning assumptions, not measured savings. InferCrane validates the real workload before recommending a route.

Roadmap

Open models today. Universal inference tomorrow.

Start with one API. Expand into a system that continuously improves execution across models, runtimes, kernels, clouds, and chips.

  1. 01

    Now · founding access

    Open models without infrastructure complexity

    Plan and deploy open models in infrastructure you control, connect an existing endpoint, or evaluate an open alternative to the OpenAI or Claude workload you already trust.

    What it unlocks

    • One compatible API
    • Open model evaluation
    • Managed capacity only when qualified
    • Existing endpoint adoption
    • Quality, privacy, latency, and cost checks
  2. 02

    Next · optimized deployments

    Performance matched to your workload

    InferCrane selects and tunes models, runtimes, decoding, kernels, and hardware around your requirements, then produces a portable deployment for managed, dedicated, or your cloud.

    What it unlocks

    • Candidate search
    • Runtime and decoding optimization
    • Kernel optimization where it matters
    • Portable execution recipes
  3. 03

    Then · continuous operations

    Inference that improves with change

    Keep your application stable while InferCrane monitors performance, requalifies new options, manages releases, and recovers from regressions.

    What it unlocks

    • Evidence-aware routing
    • Autoscaling and observability
    • Release and rollback controls
    • Drift detection and requalification
  4. 04

    North star · universal execution

    Write once. Run everywhere.

    Integrate one stable InferCrane API. Run supported open models across clouds, runtimes, and chips without rewriting your application. InferCrane finds, qualifies, and operates the best execution path as the ecosystem changes.

    What it unlocks

    • Portable model execution
    • Cross-cloud and cross-chip scheduling
    • Generated kernels
    • Programmable heterogeneous inference
Use InferCrane today

Start with one API.

Deploy an open model in your infrastructure, bring an existing endpoint, or start with the OpenAI or Claude workload you already trust.

Open-model evaluation

You do not need to choose an open model first.

Bring representative OpenAI or Claude traffic and the requirements that matter to your application. InferCrane keeps that API as the control, tests a bounded set of open-model execution paths, and recommends a move only when the evidence passes.

  • Your current API stays the pinned baseline
  • Quality and tool behavior are checked on your examples
  • Privacy, latency, and total cost are evaluated separately
Model APIs

Model API catalog, guarded availability.

The catalog and usage controls are live. Managed routes remain unavailable until a supplier-backed offer has current availability, pricing, funding, and reconciliation evidence.

  • OpenAI-compatible endpoint
  • Prepaid credits, no subscription
  • Usage limits + spending caps
  • Deployable and callable states stay distinct
  • Callable only with current supplier evidence
  • Usage and billing visibility
Dedicated & optimized

Your workload, a better build.

For your own infrastructure on AWS, GCP, or Kubernetes: dedicated capacity, custom weights, hardware migration, and runtime + kernel optimization where the evidence justifies it.

  • Custom deployments
  • Customer-owned compute
  • Candidate hardware qualified per workload
  • Runtime optimization
  • Kernel work where justified
  • Private infrastructure
Sandboxes · powered by Brezel

Give every agent its own computer.

Run coding agents, evaluations, and long jobs away from your laptop and credentials. Keep the workspace, put compute on standby when idle, and resume through one small interface.

  • Firecracker isolation on dedicated hosts
  • Durable /workspace across standby and resume
  • Commands, files, previews, and lifecycle evidence
  • Scoped model access without long-lived keys in the sandbox
FAQ

Before you move traffic.

What InferCrane changes, what it proves, and who controls deployment.