Run the model
Compile a reviewed serving plan for infrastructure you control, then deploy through the CLI, API, SDK, or Terraform.
Run InferCrane →Hugging Face to stable API
Start with a Hugging Face repository. InferCrane turns it into an inspectable serving plan before anything runs.
Plan a model →Choose a current repository or paste the exact Hugging Face ID.
Discovery metadata only. Your plan determines runtime, GPU fit, cost, and measured performance.
One endpoint contract
Compile a reviewed serving plan for infrastructure you control, then deploy through the CLI, API, SDK, or Terraform.
Run InferCrane →Put an existing OpenAI-compatible endpoint behind the same model identity, privacy acknowledgement, request limit, and cost ceiling.
Review the connection flow →A callable route appears only when supplier availability, public pricing, funding, and usage reconciliation are current. Otherwise the catalog remains fail-closed.
Check current availability →Exact model identity
Submit another Hugging Face repository, an immutable custom OCI workload, or an endpoint your team already operates. Compatibility is checked in the plan before any infrastructure mutation.
infercrane modelsinfercrane workload init ./my-model --model organization/repositoryinfercrane workload plan ./my-model