Hugging Face to stable API

Choose a model. Get a production endpoint.

Start with a Hugging Face repository. InferCrane turns it into an inspectable serving plan before anything runs.

Plan a model →
MODEL TO ENDPOINTModel identity → serving plan → stable APISimple defaults first. Advanced controls remain editable.

Find a model

Choose a current repository or paste the exact Hugging Face ID.

Hugging Face snapshot · 7/7 feeds · Oct 4, 2026

Discovery metadata only. Your plan determines runtime, GPU fit, cost, and measured performance.

ModelTaskUpdatedSignalsAction
text generationSep 1, 2026705.8K downloads · 5.1K likes
text generationSep 4, 20261.4M downloads · 2.1K likes
image text to textSep 7, 20265.4M downloads · 2.7K likes
image text to textSep 2, 20261.2M downloads · 11.6K likes
image text to textMay 19, 2026472.8K downloads · 1.6K likes
text generationAug 1, 20264.5M downloads · 4K likes

One endpoint contract

Choose where inference runs without coupling the app.

AVAILABLE NOW · OPEN SOURCE

Run the model

Compile a reviewed serving plan for infrastructure you control, then deploy through the CLI, API, SDK, or Terraform.

Run InferCrane →
AVAILABLE NOW · YOUR INFRASTRUCTURE

Connect compatible capacity

Put an existing OpenAI-compatible endpoint behind the same model identity, privacy acknowledgement, request limit, and cost ceiling.

Review the connection flow →
QUALIFIED OFFERS ONLY · MANAGED MODEL APIS

Use qualified managed capacity

A callable route appears only when supplier availability, public pricing, funding, and usage reconciliation are current. Otherwise the catalog remains fail-closed.

Check current availability →

Exact model identity

The catalog is a shortcut, not an allowlist.

Submit another Hugging Face repository, an immutable custom OCI workload, or an endpoint your team already operates. Compatibility is checked in the plan before any infrastructure mutation.

Plan from the CLI
bash
infercrane modelsinfercrane workload init ./my-model   --model organization/repositoryinfercrane workload plan ./my-model