Llama 3.1 405B Instruct

Meta · text · Open-source · GPU Tier S

liveContext 128K
Requests 30d+14.2%
1.9M
Tokens 30d+18.4%
5.4B
Revenue 30d+20.2%
$7.2K
Gross margin+1.2%
40%
Revenue − node payouts
Active users+9.4%
310
Active nodes+3.4%
2,419
Input $/1M
$1.2
Output $/1M
$1.8
p50 latency
720ms
p95 latency
2.1s
Error rate
2.3%

Usage over time (30d)

Pricing simulator

Elasticity-based what-if · Llama 3.1 405B Instruct

Projected revenue 30d
$7.2K
0.0%
Projected requests
1926k
0.0%
Projected margin
58%

Model assumes elasticity of -1.2 on total price. Replace with regression-derived elasticity per model when telemetry is wired.

p95 latency over time (30d)

Price comparison

$/1M tokens

FAR AI
Input
$1.2
Output
$1.8
Together AI
In: $3.5 · Out: $3.5
FAR saves
57%
Fireworks AI
In: $3 · Out: $3
FAR saves
50%

Revenue economics

How this model generates revenue

Developer pays
$1.2 in + $1.8 out per 1M tokens
Node operator earns
52% of revenue distributed as payout · Tier S
FAR AI gross margin
40%