HuggingFace
Description
HuggingFace operates two inference modes that the gateway supports through the same huggingface provider integration:
- Serverless inference API — the gateway routes to
https://api-inference.huggingface.co/models/<MODEL>/v1/chat/completions. The model name is taken from the request body and embedded in the URL. - Dedicated endpoint — when the gateway has
hf_endpointset in its configuration, the gateway routes to<hf_endpoint>/v1/chat/completionsinstead.
In both modes, the request and response wire format is OpenAI-compatible.
| Feature | Limitation |
|---|---|
| Chat completions | — |
| Streaming responses (SSE) | — |
| Tool use | Depends on the underlying model. |
| Vision input | Depends on the underlying model. |
| Embeddings | Not exposed by the gateway. |
Required gateway configuration
| Key | Required | Description |
|---|---|---|
hf_endpoint |
only for dedicated endpoints | The base URL of the dedicated endpoint, without trailing path. When unset, the gateway uses the serverless URL. |
BYOK key format
The BYOK value for the huggingface provider is a HuggingFace access token. The gateway sends the token in the Authorization: Bearer <KEY> header. The token must have the Inference scope.
Adding the key
The procedure for storing a BYOK key is the same for every provider. See Provider keys (BYOK) for the steps. Select huggingface in the Provider drop-down list and paste the access token in the API Key field.