Llama 3.1 70B Instruct

Meta · text · Open-source · GPU Tier A

liveContext 128K
Requests 30d+14.2%
2.9M
Tokens 30d+18.4%
6.6B
Revenue 30d+20.2%
$2.6K
Gross margin+1.2%
48%
Revenue − node payouts
Active users+9.4%
444
Active nodes+3.4%
3,589
Input $/1M
$0.35
Output $/1M
$0.55
p50 latency
420ms
p95 latency
1.3s
Error rate
1.5%

Usage over time (30d)

Pricing simulator

Elasticity-based what-if · Llama 3.1 70B Instruct

Projected revenue 30d
$2.6K
0.0%
Projected requests
2862k
0.0%
Projected margin
58%

Model assumes elasticity of -1.2 on total price. Replace with regression-derived elasticity per model when telemetry is wired.

p95 latency over time (30d)

Price comparison

$/1M tokens

FAR AI
Input
$0.35
Output
$0.55
Together AI
In: $0.88 · Out: $0.88
FAR saves
49%
Groq
In: $0.59 · Out: $0.79
FAR saves
35%
Fireworks AI
In: $0.9 · Out: $0.9
FAR saves
50%

Revenue economics

How this model generates revenue

Developer pays
$0.35 in + $0.55 out per 1M tokens
Node operator earns
44% of revenue distributed as payout · Tier A
FAR AI gross margin
48%