Integrations
How do I integrate a Hugging Face account?
How do I integrate a Hugging Face account?
- Log in to Hugging Face, then navigate to Access Tokens.
- Create a new token. You can use a fine-grained token. In this case, make sure the token has view permission for the repository you’d like to use.
- Integrate the key in Friendli Suite > Personal Settings > Integrations.
If you revoke / invalidate the key, you will have to update the key to avoid disrupting ongoing deployments, or to launch a new inference deployment.
Using a 3rd-Party Model
How can I use a Hugging Face repository as a model?
How can I use a Hugging Face repository as a model?

- Use the repository id of the model. You can select the entry from the list of autocompleted model repositories.
- You can select a specific branch, or manually enter a commit hash.
Format Requirements
What are the format requirements for a model?
What are the format requirements for a model?
- A model should be in safetensors format.
- The model should NOT be nested inside another directory.
- Including other arbitrary files (that are not in the list) is totally fine. However, those files will not be downloaded nor used.
What are the format requirements for a dataset?
What are the format requirements for a dataset?
The dataset should satisfy the following conditions:
- The dataset must contain a column named “messages”.
- Each row in the “messages” column should be compatible with the chat template of the base model.
For example,
tokenizer_config.jsonofmistralai/Mistral-7B-Instruct-v0.2is a template that repeats the messages of a user and an assistant. Concretely, each row in the “messages” field should follow a format like:[{"role": "user", "content": "The 1st user's message"}, {"role": "assistant", "content": "The 1st assistant's message"}]. In this case,HuggingFaceH4/ultrachat_200kis a dataset that is compatible with the chat template.
Troubleshooting
Inference Request Errors
Common error codes for inference requests
Common error codes for inference requests
Below is a table of common error codes you might encounter when making inference-related API requests.
Quick Checklist Before Retrying
- Verify the endpoint URL,
endpoint_id, and (if applicable)X-Friendli-Teamheader - Include the
Authorizationheader with a valid key - Confirm the target deployment exists, is healthy, and is not deleted
- Validate request JSON and required fields; reduce
max_tokensif needed - Check rate limits; add retry with backoff when receiving
429
Model Selection Errors
You don't have access to this gated model
You don't have access to this gated model

The repository / artifact is invalid
The repository / artifact is invalid


The architecture is not supported
The architecture is not supported

Endpoint Lifecycle
Why was my endpoint suddenly terminated?
Why was my endpoint suddenly terminated?
Endpoints that remain in a sleep state for 48 hours are automatically terminated.
- When
min_replicas = 0, the endpoint enters a sleep state after the cooldown period if no requests are received. - A notification is sent after 24 hours of sleep, and the endpoint is terminated after another 24 hours if not reactivated.