Inference providers

Sell inference through FastMetal

Your OpenAI-compatible endpoint becomes a model in the FastMetal catalog. Japanese developers buy it in yen; FastMetal handles billing, consumption tax and qualified invoices, and pays you monthly against a token statement you can check yourself.

Currently serving 53 models from 5 upstream hosts.

We prioritize models hosted in Japan and proprietary models that are strong in Japanese. Applications are reviewed on a rolling basis and not every application is accepted.

Want to build FastMetal into your own product instead? See the partner and OEM page.

How it works

One API surface

Every customer reaches your models the same way: one key, /v1/chat/completions, the chat UI and the MCP server. You integrate once and every FastMetal surface routes to you.

Yen billing, tax and invoices

Customers pay in yen by card. FastMetal is a registered qualified invoice issuer in Japan, so consumption tax and 適格請求書 are handled on our side, not yours.

Japan-hosted, said out loud

Models that run in Japan are labelled as such in the catalog and on their model page, and sold to customers who specifically want inference inside Japan.

Monthly settlement you can verify

Your provider console shows requests and tokens per model per day, exportable as CSV. You invoice FastMetal against that statement.

Technical requirements

The technical review checks these with real calls through the gateway.

  • An OpenAI-compatible /chat/completions endpoint with streaming, returning usage (token counts) on both streaming and non-streaming responses. Usage is the billing basis on both sides.
  • A /models endpoint, or a written rate card, with price, context length and max output tokens per model.
  • A published privacy policy and a stated position on whether prompts are used for training and how long they are retained. We record it with your listing and pass it on to customers who ask.
  • Where inference physically runs, in writing. Japan-hosted models are labelled; the label is what we sell.

Declared, not required

Tell us which of these your endpoint supports; each is verified before it is advertised.

  • Tool calling (function calling) and structured outputs
  • Image input
  • Prompt caching, with cache read and write pricing

Process

  1. 1

    Apply

    Tell us about your infrastructure, endpoints, models, pricing and data policy. We reply by email.

  2. 2

    Technical review

    We make real calls through the gateway: compatibility, streaming, usage reporting, tool calling and image input where claimed, and pricing reconciled against your rate card.

  3. 3

    Staging

    Your models are added to a hidden staging configuration, run through the integration tests, and given real example outputs for the comparison viewer and their model pages.

  4. 4

    Go live

    Your models go into the production catalog with their own pages, an announcement post on the blog and on X, and traffic from every FastMetal surface.

Questions providers ask

How is pricing decided?
You quote a wholesale rate (input and output per million tokens); FastMetal sets the customer-facing price in yen and publishes it on the model and pricing pages. Wholesale rates and terms are agreed per provider and are not published.
How do we get paid?
Monthly. The provider console shows requests and tokens per model per day, exportable as CSV; you invoice FastMetal against that statement. Payment terms are agreed when we sign.
What do we have to support?
An OpenAI-compatible /chat/completions endpoint with streaming that returns usage on both streaming and non-streaming responses, and either a /models endpoint or a written rate card giving price, context length and max output tokens per model. You also need a published privacy policy and a stated position on training and retention of prompts.
Which providers are prioritized?
Models hosted in Japan, and proprietary models strong in Japanese. Japan-hosted models are labelled as such in the catalog and sold to customers who specifically want inference inside Japan.
How are prompts and data handled?
Your training and retention policy is recorded with your listing and passed on to customers who ask; showing it on the model page itself is planned. For direct API requests, FastMetal stores nothing beyond what billing needs — token counts, model name and time — never the prompt. Requests made through FastMetal’s own chat UI are stored as that customer’s conversation history on FastMetal.
Can companies outside Japan apply?
Yes, although Japan-hosted models are prioritized. For inference outside Japan, what matters is whether the model is distinctive or strong in Japanese.
How is this different from OpenRouter?
FastMetal sells to Japanese customers in yen, handles consumption tax and issues qualified invoices, and labels Japan-hosted models as such. The network is smaller than OpenRouter’s; in return each listed model gets individual attention.

Ready to apply?

The form takes about ten minutes. We reply to every application by email, usually within a few business days.

Apply to become a provider