Production inferencewithout the platform engineering.
Deploy a model across infrastructure you control and expose one stable API. InferCrane handles the durable lifecycle, autoscaling, monitoring, and evidence-gated releases behind it. Already running inference? Connect it without migrating first.
$ infercrane deploy mistralai/Mistral-7B-Instruct-v0.3Creating support-production
desired state / persistedinfercrane operation watch op_7f3aBuilt on proven inference infrastructure—not a replacement for it
Start with the outcome
From a model to one production endpoint.
Build a new inference service, govern a model API, or adopt an existing stack. Every path converges on the same stable endpoint, durable lifecycle, and evidence model.
Deploy a production endpoint
Start from a reviewed recipe or explicit serving plan and let one durable operation own the lifecycle.
✓ Model → stable endpointSee deployment path →Create one governed model API
Give the application a stable model name with explicit privacy, traffic, and cost controls.
✓ Credentials stay server-sideSee API composition →Connect without migration
Discover vLLM, SGLang, LiteLLM, or another OpenAI-compatible endpoint read-only.
✓ Observe → route → manageSee adoption path →Verified Models
Start quickly. Know exactly what is—and is not—verified.
Reviewed immutable configurations remove first-deploy guesswork. Performance, price, provider capacity, and GPU fit remain evidence-backed qualification decisions.
BGE-M3
BAAI/bge-m3embeddings · retrieval · RAG
MIT · configuration verifiedQwenQwen2.5 Coder 7B Instruct
Qwen/Qwen2.5-Coder-7B-Instructcoding · chat
Apache-2.0 · configuration verifiedGoogleGemma 3 4B IT
google/gemma-3-4b-itvision · chat · documents
Gemma · configuration verifiedFrom model to production
Deploy, serve, operate, and improve—in one loop.
A guided tour of the private-preview console using representative, contract-faithful evidence. No provider performance or production-readiness claim is implied.
Start from a reviewed model identity—not an unexplained preset.
Browse immutable model revisions, licenses, protocols, and runtime capabilities. Starting configurations are clearly separated from measured benchmark evidence, and any explicit Hugging Face model remains available.
✓ CONFIGURATION VERIFIED · PERFORMANCE UNCLAIMEDconsole.infercrane / chooseSIMULATED EVIDENCE
Operational outcomes
Spend less guesswork on cost, latency, and releases.
InferCrane does not promise a magic percentage. It makes each optimization measurable, reviewable, and reversible against the workload evidence you actually have.
Know what inference costs—and where the waste is.
Attribute sourced GPU and model-API cost, bound external fallback, and compare serving plans without turning missing prices into savings claims.
cost / requestcost / 1M tokensidle capacityfallback spendFind latency pressure before a candidate reaches production.
Correlate TTFT, total latency, queueing, throughput, cold starts, and capacity. Release Guard rejects regressions using persisted policy and evidence.
TTFT p95latency p95queue p95output tok/sRemove terminal babysitting and rollout scripts.
One durable operation owns provisioning, readiness, routing, draining, and cleanup. Close the terminal and reconnect without corrupting the deployment.
durable operationstable endpointsafe draindeterministic rollbackChoose with evidence
API, self-hosted, or hybrid is a workload decision.
Start with the operating model that fits today. InferCrane keeps one application endpoint while benchmark, replay, cost, privacy, and capacity evidence inform the next serving plan.
recommendations advise · humans approve · Release Guard verifiesThe missing operational layer
Running a model is one command.Operating it safely is a system.
InferCrane owns the lifecycle and evidence around inference. Proven infrastructure projects keep doing the jobs they already do well.
Build from zero
From model to durable production endpoint.
Start with a model or immutable OCI workload. InferCrane turns the serving plan into a durable operation, publishes one application-facing endpoint, and keeps deployment, scaling, monitoring, and release state attached to it.
infercrane deploy mistralai/Mistral-7B-Instruct-v0.3 \ --name support-productioninfercrane status support-production --watchinfercrane connect https://vllm.internal/v1 \ --as coder-production --type vllminfercrane request inspect req_01J...One endpoint · replaceable stack
Use the infrastructure specialist for each job.
InferCrane does not become your provider catalog, vector database, agent framework, sandbox runtime, training scheduler, or workflow engine. It connects those systems through explicit contracts and keeps the production inference decision in one place.
Start with a model API. Keep the exit door open.
Put one stable, budgeted endpoint in front of OpenRouter or another OpenAI-compatible provider. Add self-hosted capacity later without changing application code.
✓ No traffic until consent and hard budgets are explicitinfercrane provider connect openrouter-main \ --model openai/gpt-4.1-mini --from-env OPENROUTER_API_KEYRelease Guard
A healthy container can still ship a worse model.
Release Guard compares the active and candidate serving plans using trustworthy measurements. Decisions are deterministic, persisted, and available for audit.
- Readiness and runtime failures
- TTFT, latency, throughput, and errors
- Benchmark, replay, and signed evaluator evidence
Operational evidence
Know why. Not just what.
Follow a request across its logical endpoint, revision, and replica. Explain degraded deployments, scaling, rollouts, and cold starts from recorded facts—not an LLM guess.
infercrane doctor coder-productionModular by contract
Bring your clouds. Bring your runtimes. Keep one operating model.
Core state never depends on a provider or engine. Adapters translate a serving plan into infrastructure, runtime, and traffic behavior—and report qualification evidence separately.
model="coder-production"stableThe private-preview platform
The inference operating loop, covered.
InferCrane owns the durable control and evidence plane. Availability and qualification vary by provider and runtime; unknown evidence is never converted into a claim.
Start from your model or your running endpoint.
- Reviewed model recipes
- Existing endpoint discovery
- vLLM · SGLang · custom OCI
- Immutable artifact identity
Keep application traffic stable while capacity changes.
- Stable logical endpoints
- Elastic and serverless modes
- Autoscaling · queueing · quotas
- Governed external fallback
Turn runtime signals into operational evidence.
- OpenTelemetry · DCGM · OpenCost
- Request Inspector · Doctor
- Cold-start and scaling timelines
- Alerts · cost attribution
Change serving plans without relying on hope.
- Immutable revisions and rollback
- Release Guard
- AIPerf benchmark · workload replay
- Signed Inference Passport
Compare options using labeled evidence.
- Inference Lab
- Capacity intelligence
- Artifact cache and prefetch evidence
- Advisory FinOps recommendations
Keep specialist systems behind explicit boundaries.
- LiteLLM and model APIs
- Training artifact handoff
- Scoped sandbox access
- Agents · RAG · external workflows
Private preview · invitation only
Operate your first endpoint with us.
We are inviting a small number of teams running—or preparing to run—production inference. Join the list for preview access and launch updates.