Llama 3.1 8B Instruct

Meta · text · Open-source · GPU Tier B

liveContext 128K
Requests 30d+14.2%
2.3M
Tokens 30d+18.4%
5.9B
Revenue 30d+20.2%
$335.2
Gross margin+1.2%
54%
Revenue − node payouts
Active users+9.4%
528
Active nodes+3.4%
2,888
Input $/1M
$0.05
Output $/1M
$0.08
p50 latency
280ms
p95 latency
720ms
Error rate
1.3%

Usage over time (30d)

Pricing simulator

Elasticity-based what-if · Llama 3.1 8B Instruct

Projected revenue 30d
$335.2
+0.0%
Projected requests
2301k
0.0%
Projected margin
58%

Model assumes elasticity of -1.2 on total price. Replace with regression-derived elasticity per model when telemetry is wired.

p95 latency over time (30d)

Price comparison

$/1M tokens

FAR AI
Input
$0.05
Output
$0.08
Together AI
In: $0.18 · Out: $0.18
FAR saves
64%
Groq
In: $0.05 · Out: $0.08
+-0%
Fireworks AI
In: $0.2 · Out: $0.2
FAR saves
68%

Revenue economics

How this model generates revenue

Developer pays
$0.05 in + $0.08 out per 1M tokens
Node operator earns
38% of revenue distributed as payout · Tier B
FAR AI gross margin
54%