01 /TOOLS / groq
Groq
shipUltra-fast inference on custom LPU hardware — open models at 500+ tok/s.
journal entry · observed by @whysanesanders · verified 2026-08-15
Speed that is a feature, not a benchmark: sub-second responses change what UX you can build — voice agents, autocomplete, real-time tools. The free tier is generous enough to build on.
Known limitations
- Model catalog limited to what fits their hardware.
- Context windows capped below the biggest hosted rivals.
- Free-tier throughput limits are tight for production.
Facts
- pricing
- freemium
- price note
- free tier; from $0.05/1M tokens (small models)
- free tier
- yes
- open source
- no
- api
- yes
- self-host
- no
- category
- dev-infra
Models used
Receipts
Same sector — dev-infra
Fireworks AI
Fast inference platform for open models — serverless and on-demand GPUs.
dev-infra
freemium · free credits; pay-per-token from ~$0.10/1M (small models)
APIFREE TIER
Hugging Face
The GitHub of AI — models, datasets, Spaces and inference endpoints.
dev-infra
freemium · free; PRO $9/mo; inference pay-as-you-go
APIFREE TIER
OpenRouterfeatured
One API for every major model — unified billing, fallbacks and routing.
dev-infra
freemium · pay-per-token pass-through plus a small fee
APIFREE TIER