Cost estimate

AI cost calculator

Enter your monthly bill, your token volume or your headcount to see what the same work costs in yen on another model. Every model FastMetal serves is compared, not only the open-weight ones.

Prices are FastMetal's listed rates as of October 1, 2026 (per 1M tokens, excluding consumption tax).

Enter your current monthly bill and pick the main kind of use. Token volume is worked back from these.

Worked back at 4 input tokens per output token, with 0% of input read from cache.

FastMetal's price: input ¥357.4 / output ¥1,787 / cache read ¥35.74

Not all work can move to a cheaper model. Keep the hard work where it is and choose the share you can move.

Estimated saving

¥496,866a month

That is 50% of the current monthly cost, or ¥5,962,395 a year.

Monthly cost now and after the move
Monthly cost now¥1,000,000
Monthly cost after¥503,134
  • Kept on Claude Sonnet 5.5: ¥500,000
  • Moved to Llama 3.1 8B Instruct: ¥3,134

Assumed volume: 1,243.5 input and 310.9 output (millions of tokens a month), 0% of input read from cache

Models that cost less at the same volume

69 models cost less a month than Claude Sonnet 5.5. Select a row to use it in the estimate above.

The baseline model has no LMArena score, so the list isn't filtered by score.

ModelLMArena scoreInputOutputMonthly cost if all work movesvs. now
1211¥3.36¥6.72¥6,267−99%
—¥8.4¥8.4¥13,057−99%
1318¥5.04¥23.52¥13,580−99%
—¥7.15¥26.81¥17,226−98%
—¥10.08¥30.24¥21,936−98%
—¥8.94¥35.74¥22,228−98%
1436¥15.12¥30.24¥28,204−97%
1167¥8.94¥58.97¥29,450−97%

The LMArena score is the value from LMArena's public text leaderboard, overall category (CC BY 4.0), the same one shown on each model page. A close score doesn't show that two models give the same result on your own work. Claude Sonnet 5.5 scores —.

Two ways to keep the cost down

Match the model to the job

Summaries, classification and routine support replies rarely need the top model. Moving that work to a lower-priced model cuts the cost of each call. Compare the results on your own data before you move it.

Put a ceiling on usage

A lower price doesn't help if usage keeps growing. FastMetal is prepaid: once the balance you bought is used up, requests are refused, and there is no month-end bill in arrears. Balances are held per API key, so a key can serve as the limit for one use.

How the estimate is calculated

  • Monthly cost is input tokens times the input price plus output tokens times the output price. Input read from cache uses each model's cache-read price, and a model without one is charged its normal input price. The premium for cache writes is left out on both sides of the comparison.
  • When token volume is worked back from a bill, the price used is FastMetal's price for that model. It can differ from what you pay the vendor directly.
  • The same text is a different number of tokens on each model. This estimate assumes the token count doesn't change.
  • The per-person volumes in the headcount estimate are illustrations, not measurements of real users. The agent figure is set to land near the average reported for Uber's engineers (US$150 to 250 a month).
  • Models with close LMArena scores still differ in how well they hold up over long tasks and how accurately they call tools. That is why the share moved starts at 50%.

Frequently asked questions

How much cheaper is an open-weight model?

It depends on the model and on how you use it. The table on this page works out the monthly cost per model at the same volume, including the split between input, output and cache reads. A model with a lower price but no cache-read discount may not be cheaper for agent-style use.

Can I switch the model behind a coding agent?

Only testing on your real development work can tell you. Models with close leaderboard scores still differ in how well they hold up over long tasks and how accurately they call tools. That is why the calculator lets you choose the share you move and doesn't assume all of it.

Can it estimate the cost of giving every employee AI?

Use the "From headcount" tab: enter how many people mostly chat and how many use agents, and it gives a monthly cost and a per-person figure. FastMetal provides an API and a chat screen, not central management of employee accounts. The estimate assumes your internal tools call the API.

When are the prices from?

They are the prices FastMetal lists when you open this page, excluding consumption tax. The pricing page has every model's price.