Yasin Toy, founder of InferCrane · October 8, 2026Open weights · Apache-2.0 metadata

Introducing Commerce-1

An open
decision model
for commerce agents.

Today we're releasing Commerce-1: a self-hosted 27B model that turns structured state into typed Choice, Noul, and Score distributions—without generating an explanation first.

Public index61.69Decision Index 0.3 public component
Coverage140,178 / 140,178scoreable requests completed
Output0 tokenstyped answers, not generated prose
StatusPublic run completeFull Score and rank pending

Agentic commerce fails at the handoff from conversation to action.

An agent can name the right product and still use the wrong store location. It can recommend a convincing item that is already sold out, or offer an alternative that quietly breaks the customer's budget, dietary constraint, or delivery window. These are not writing failures. They are decision failures at the boundary between language and live commerce state.

A generative model should still browse, reason, and plan. But the moment before an action is narrower: present one eligible offer, call one allowed tool, ask for a missing fact, route to a person, or stop. Generating another paragraph solves a larger problem than the application has and leaves code to infer the decision.

Commerce-1 is deliberately narrower. It scores supplied actions directly and returns a typed answer with a probability for every option. The application then applies fresh inventory and location truth, budget rules, policy gates, authorization, and its own audit trail. The model proposes; code remains in charge.

Three small primitives cover a surprising amount of work.

Ask multiple ordered questions about one structured state in one API request.

01

Choice

Which eligible offer should the agent present?

selected option + probability for every option
02

Noul

Does this action need human review?

boolean answer + P(true)
03

Score

How risky is this checkout state?

expected ordinal score + full distribution

The qualified runtime scores each question separately with question microbatch 1. An API request may carry multiple questions; this is not a claim that every question is answered in one model pass.

A direct path from state to distributions.

Full-attention layers receive noncausal masks. Recurrent linear-attention layers retain their original recurrence.

Commerce-1 uses last-position pooling and a dedicated decision readout to score supplied options. Deterministic application code remains responsible for hard constraints and authorization.

We optimized the first release for capability, not for the smallest box.

Many decision models start small. That is attractive for latency and local use, but commerce state is rarely clean: policies conflict, products are described inconsistently, and an action can depend on facts spread across a long input. We wanted to test the decision interface without making small-model capacity the first bottleneck.

The tradeoff is explicit. Commerce-1's BF16 weights are about 54 GiB, and the release path is qualified on one NVIDIA H200—not on a laptop, CPU, or 24/48 GiB GPU. A smaller distilled model is a future optimization. This release establishes the higher-capability, inspectable baseline first.

A result is useful only when its boundary travels with it.

Decision Index 0.3Public component receipt
61.69public index
Raw public score
71.28
Scoreable requests
140,178 / 140,178
Unsupported
0
Errors
0
Full Score
Pending
Official rank
Not claimed
Inspect immutable results
Capability profileChance-corrected public area scores
Tools & automation77.93
Language67.32
Retrieval & classification60.06
Knowledge & reasoning51.56
Arts & human taste46.93

These are public-suite area scores, not production success rates.

The self-hosted 0.3 public run completed every scoreable request. It remains one input to the benchmark's Full Score; private components, held-out-answer checks, official latency, and rank remain with the maintainers. The immutable release was initially gated on a separate complete 0.2.1 reproduction at 60.99.

One exact package. One exact path.

01 / HardwareNVIDIA H2001 GPU · BF16 · network disabled
p50340.146 mscomplete Choice + Noul + Score fixture
p95349.026 mscomplete three-decision fixture
Throughput8.836 decisions/ssame locked fixture · microbatch 1
Parity0.0 max difference9 native/direct + 9 release/evaluated comparisons

Twelve measured iterations after two warmups on BF16, CUDA 13.0, and PyTorch 2.14.0. This is content-bound package qualification—not one-decision HTTP latency, the benchmark's official latency, or a promise for other hardware and traffic.

One customer. Four baskets. Make the agent's call.

You decide what a shopping agent should recommend for four friends, one strict nut allergy, thirty minutes, and an €18 budget. Pick one fictional basket; then compare your call with Commerce-1 and a separate deterministic checkout rule.

Take the 20-second test
Northstar MarketCommerce-1 receiptprecomputed typed output
Tomato lentil spaghetti€8.45
BudgetPASS
VegetarianPASS
Strict nut-freePASS
Choice probability82.99%

Synthetic store and prices. Real precomputed Commerce-1 output. Deterministic code, not the model, checks hard constraints. The user owns checkout.

The native API accepts state plus ordered questions.

Run the immutable model and source revisions, start the offline server, then send a bounded decision. The response contains typed answers and ordered probabilities.

curl -s http://127.0.0.1:8000/v1/systemone \
  -H 'content-type: application/json' \
  --data '{
    "model": "infercrane/commerce-1",
    "state": {"cart_total": 117, "approved_total": 94},
    "questions": {
      "route": {
        "type": "choice",
        "instructions": "Choose the safest next action.",
        "criteria": {
          "continue": "Continue.",
          "review": "Ask for review.",
          "block": "Stop."
        }
      }
    }
  }'

Commerce-1 is not a live-fact retriever or payment authority. Recalibrate on your own traffic, keep consequential actions behind explicit controls, and treat probability as evidence—not permission.

A component for bounded judgment—not an autonomous buyer.

Designed for

  • Proceed, review, or block gates
  • Tool and workflow routing
  • Product, fulfillment, and support action selection
  • Risk and priority scoring

Not designed as

  • A chat model or live-fact retriever
  • A payment processor or fraud oracle
  • Sole authority for consequential decisions
  • A qualified vision, audio, CPU, mobile, or laptop model