Skip to content
GridRouterhome

Search

Search providers, capabilities and pages

GroqCloud API

groq.com

Verification pendingResearched, not yet routed
Low-latency inference API for open models (OpenAI-compatible). Buying direct: Free plan; Developer plan pay per token. Listed .
Capabilities
0
From
n/a
Calls, 30 days
0
Ranked first
0 of 0

Overview

Researched from the vendor's own pages; every fact links to its source and the day we checked it.

GroqCloud serves open-weight LLMs, speech-to-text and TTS on LPU hardware through an OpenAI-compatible API. Rate limits are per organization in RPM, RPD, TPM, TPD (and audio seconds); Free and Developer plans, with headers x-ratelimit-* and retry-after on 429.

Company

Legal name
Not confirmed
Headquarters
Not confirmed
Founded
Not confirmed
Employees
Not confirmed
Ownership
Not confirmed
Kind
Infrastructure
Website
Not confirmed

API

Verified
Access
Self-serve API
Style
REST
Auth
Not confirmed
Base URL
Not confirmed
Features
Bulk —Async —Webhooks —MCP —Sandbox —
Rate limit
Per org per model: e.g. Free plan gpt-oss-120b 30 RPM, 1K RPD, 8K TPM, 200K TPD

Pricing

Model
pay as you go
List prices
Contact sales
Free tier
Free plan with base rate limits

Compliance

SOC 2
Not confirmed
GDPR
Not confirmed
CCPA
Not confirmed
DPA
Not confirmed
Resale / storage
Unknown until a written agreement says otherwise

Relationships

Integrations
Not confirmed
Alternatives
None named by the vendor
Sources (2)

Placement in the Grid

Where this vendor sits on the GridRouter Grid in each category it serves. No vendor pays for placement.

Placement in the GridRouter Grid by category
CategoryPlacementExecuteCoverageCallsReport
AI UtilitiesNot yet rated––0View Grid

API

The vendor's OpenAPI document as stored in the catalog, its operations, and the rate limits it publishes.

No spec

Rate limits

Published rate limits for groq
LimitScopePlanEndpointBurst / concurrency
30 / minuteaccountFreeopenai/gpt-oss-120b— / —
30 / minuteaccountFreeopenai/gpt-oss-20b— / —
30 / minuteaccountFreemeta-llama/llama-prompt-guard-2-86m— / —

Headers: x-ratelimit-limit-requestsx-ratelimit-limit-tokensx-ratelimit-remaining-requestsx-ratelimit-remaining-tokensx-ratelimit-reset-requestsx-ratelimit-reset-tokensretry-after. GridRouter never paces above these and backs off on 429, honoring Retry-After. Source, checked Sep 29, 2026

Compare GroqCloud