Llama 3.1 70B Instruct
Meta · text · Open-source · GPU Tier A
liveContext 128K
Requests 30d+14.2%
2.9M
Tokens 30d+18.4%
6.6B
Revenue 30d+20.2%
$2.6K
Gross margin+1.2%
48%
Revenue − node payouts
Active users+9.4%
444
Active nodes+3.4%
3,589
Input $/1M
$0.35
Output $/1M
$0.55
p50 latency
420ms
p95 latency
1.3s
Error rate
1.5%
Usage over time (30d)
Pricing simulator
Elasticity-based what-if · Llama 3.1 70B Instruct
Projected revenue 30d
$2.6K
0.0%
Projected requests
2862k
0.0%
Projected margin
58%
Model assumes elasticity of -1.2 on total price. Replace with regression-derived elasticity per model when telemetry is wired.
p95 latency over time (30d)
Price comparison
$/1M tokens
FAR AI
Input
$0.35
Output
$0.55
Together AI
In: $0.88 · Out: $0.88
FAR saves
−49%
Groq
In: $0.59 · Out: $0.79
FAR saves
−35%
Fireworks AI
In: $0.9 · Out: $0.9
FAR saves
−50%
Revenue economics
How this model generates revenue
Developer pays
$0.35 in + $0.55 out per 1M tokens
Node operator earns
44% of revenue distributed as payout · Tier A
FAR AI gross margin
48%