Llama 3.1 8B Instruct
Meta · text · Open-source · GPU Tier B
liveContext 128K
Requests 30d+14.2%
2.3M
Tokens 30d+18.4%
5.9B
Revenue 30d+20.2%
$335.2
Gross margin+1.2%
54%
Revenue − node payouts
Active users+9.4%
528
Active nodes+3.4%
2,888
Input $/1M
$0.05
Output $/1M
$0.08
p50 latency
280ms
p95 latency
720ms
Error rate
1.3%
Usage over time (30d)
Pricing simulator
Elasticity-based what-if · Llama 3.1 8B Instruct
Projected revenue 30d
$335.2
+0.0%
Projected requests
2301k
0.0%
Projected margin
58%
Model assumes elasticity of -1.2 on total price. Replace with regression-derived elasticity per model when telemetry is wired.
p95 latency over time (30d)
Price comparison
$/1M tokens
FAR AI
Input
$0.05
Output
$0.08
Together AI
In: $0.18 · Out: $0.18
FAR saves
−64%
Groq
In: $0.05 · Out: $0.08
Fireworks AI
In: $0.2 · Out: $0.2
FAR saves
−68%
Revenue economics
How this model generates revenue
Developer pays
$0.05 in + $0.08 out per 1M tokens
Node operator earns
38% of revenue distributed as payout · Tier B
FAR AI gross margin
54%