Skip to content

HuggingFace

Description

HuggingFace operates two inference modes that the gateway supports through the same huggingface provider integration:

  • Serverless inference API — the gateway routes to https://api-inference.huggingface.co/models/<MODEL>/v1/chat/completions. The model name is taken from the request body and embedded in the URL.
  • Dedicated endpoint — when the gateway has hf_endpoint set in its configuration, the gateway routes to <hf_endpoint>/v1/chat/completions instead.

In both modes, the request and response wire format is OpenAI-compatible.

Feature Limitation
Chat completions —
Streaming responses (SSE) —
Tool use Depends on the underlying model.
Vision input Depends on the underlying model.
Embeddings Not exposed by the gateway.

Required gateway configuration

Key Required Description
hf_endpoint only for dedicated endpoints The base URL of the dedicated endpoint, without trailing path. When unset, the gateway uses the serverless URL.

BYOK key format

The BYOK value for the huggingface provider is a HuggingFace access token. The gateway sends the token in the Authorization: Bearer <KEY> header. The token must have the Inference scope.


Adding the key

The procedure for storing a BYOK key is the same for every provider. See Provider keys (BYOK) for the steps. Select huggingface in the Provider drop-down list and paste the access token in the API Key field.