Meta Llama API: models and pricing

FastMetal serves 8 Meta Llama models through one OpenAI-compatible API key, billed per token from a prepaid balance. Change the base URL and key, and existing code calls them as is. The list also shows Meta Llama models we do not serve yet.

Sort by:

19 models

Muse Spark 1.3

muse-spark-1.3

Muse Spark 1.3 is a multimodal reasoning model from Meta for long-running agentic, multi-agent, and coding workflows. It is designed to keep track of information across extended tasks, work through...

Provider:meta logometa
Pricing:$1.38 in · $4.68 out / 1M tokens
Context:1.0M tokens

Muse Spark 1.3 Contributor

muse-spark-1.3-contributor
Trains on prompts

Muse Spark 1.3 Contributor is the cost-efficient contributor tier of Meta’s multimodal reasoning model for experimentation, learning, and early-stage agentic, multi-agent, and coding workflows. It is designed to track information...

Provider:meta logometa
Pricing:$0.11 in · $0.22 out / 1M tokens
Context:1.0M tokens

Muse Spark 1.2 Contributor

meta/muse-spark-1.2-contributor

Muse Spark 1.2 contributor tier is a reasoning model from Meta designed for developers who want to start building at an even lower cost. It’s meaningfully cheaper than Muse Spark...

Provider:meta logometa
Pricing:Not on FastMetal yet
Context:1.0M tokens

Muse Glimmer 30B

muse-glimmer-30b

Meta's dense, open-weight 30B multimodal model from Meta Superintelligence Labs, distilled from Muse Spark and optimized for long-horizon autonomous-agent tasks on consumer hardware.

Provider:meta logometa
Pricing:$0.3693 in · $1.59 out / 1M tokens
Context:131K tokens

Muse Spark 1.2

muse-spark-1.2

Muse Spark 1.2 is a reasoning model from Meta, designed for complex agentic tasks. It accepts text, images, video, audio, and PDF documents, returns text, and offers a 1M-token context...

Provider:meta logometa
Pricing:$1.38 in · $4.68 out / 1M tokens
Context:1.0M tokens

Muse Spark 1.1

meta/muse-spark-1.1

Muse Spark 1.1 is a multimodal reasoning model from Meta, built for agentic tasks. It accepts text, images, video, audio, and PDF documents and returns text, with a 1M-token context...

Provider:meta logometa
Pricing:Not on FastMetal yet
Context:1.0M tokens
Text Overall
#12

Llama 4 Maverick

llama-4-maverick
ZDR

Llama 4 Maverick 17B Instruct (128E) is a high-capacity multimodal language model from Meta, built on a mixture-of-experts (MoE) architecture with 128 experts and 17 billion active parameters per forward pass (400B total).

Provider:meta-llama logometa-llama
Pricing:$0.22 in · $0.88 out / 1M tokens
Context:1.0M tokens

Llama 4 Scout

meta-llama/llama-4-scout

Llama 4 Scout 17B Instruct (16E) is a mixture-of-experts (MoE) language model developed by Meta, activating 17 billion parameters out of a total of 109B.

Provider:meta-llama logometa-llama
Pricing:Not on FastMetal yet
Context:328K tokens

Llama 3.3 70B Instruct

llama-3.3-70b-instruct

The Meta Llama 3.3 multilingual large language model (LLM) is a pretrained and instruction tuned generative model in 70B (text in/text out).

Provider:meta-llama logometa-llama
Pricing:$0.1055 in · $0.3376 out / 1M tokens
Context:131K tokens
Text Overall
#261

Llama 3.3 70B Instruct (free)

meta-llama/llama-3.3-70b-instruct:free

The Meta Llama 3.3 multilingual large language model (LLM) is a pretrained and instruction tuned generative model in 70B (text in/text out).

Provider:meta-llama logometa-llama
Pricing:Not on FastMetal yet
Context:128K tokens

Llama 3.2 1B Instruct

meta-llama/llama-3.2-1b-instruct

Llama 3.2 1B is a 1-billion-parameter language model focused on efficiently performing natural language tasks, such as summarization, dialogue, and multilingual text analysis.

Provider:meta-llama logometa-llama
Pricing:Not on FastMetal yet
Context:60K tokens
Text Overall
#395

Llama 3.2 3B Instruct

llama-3.2-3b-instruct
ZDR

Llama 3.2 3B is a 3-billion-parameter multilingual large language model, optimized for advanced natural language processing tasks like dialogue generation, reasoning, and summarization.

Provider:meta-llama logometa-llama
Pricing:$0.055 in · $0.363 out / 1M tokens
Context:131K tokens
Text Overall
#366

Llama 3.2 3B Instruct (free)

meta-llama/llama-3.2-3b-instruct:free

Llama 3.2 3B is a 3-billion-parameter multilingual large language model, optimized for advanced natural language processing tasks like dialogue generation, reasoning, and summarization.

Provider:meta-llama logometa-llama
Pricing:Not on FastMetal yet
Context:131K tokens

Llama 3.1 405B (base)

meta-llama/llama-3.1-405b

Meta's latest class of model (Llama 3.1) launched with a variety of sizes & flavors. This is the base 405B pre-trained version. It has demonstrated strong performance compared to leading closed-source models in human evaluations.

Provider:meta-llama logometa-llama
Pricing:Not on FastMetal yet
Context:33K tokens

Llama 3.1 405B Instruct

meta-llama/llama-3.1-405b-instruct

The highly anticipated 400B class of Llama3 is here! Clocking in at 128k context with impressive eval scores, the Meta AI team continues to push the frontier of open-source LLMs.

Provider:meta-llama logometa-llama
Pricing:Not on FastMetal yet
Context:131K tokens

Llama 3.1 70B Instruct

meta-llama/llama-3.1-70b-instruct

Meta's latest class of model (Llama 3.1) launched with a variety of sizes & flavors. This 70B instruct-tuned version is optimized for high quality dialogue usecases.

Provider:meta-llama logometa-llama
Pricing:Not on FastMetal yet
Context:131K tokens
Text Overall
#286

Llama 3.1 8B Instruct

llama-3.1-8b-instruct

Meta's latest class of model (Llama 3.1) launched with a variety of sizes & flavors. This 8B instruct-tuned version is fast and efficient. It has demonstrated strong performance compared to...

Provider:meta-llama logometa-llama
Pricing:$0.0211 in · $0.0422 out / 1M tokens
Context:131K tokens
Text Overall
#340

Llama 3 70B Instruct

meta-llama/llama-3-70b-instruct

Meta's latest class of model (Llama 3) launched with a variety of sizes & flavors. This 70B instruct-tuned version was optimized for high quality dialogue usecases.

Provider:meta-llama logometa-llama
Pricing:Not on FastMetal yet
Context:8K tokens
Text Overall
#302

Llama 3 8B Instruct

meta-llama/llama-3-8b-instruct

Meta's latest class of model (Llama 3) launched with a variety of sizes & flavors. This 8B instruct-tuned version was optimized for high quality dialogue usecases.

Provider:meta-llama logometa-llama
Pricing:Not on FastMetal yet
Context:8K tokens
Text Overall
#333
Meta Llama API: models and pricing - FastMetal