10:24
2026-08-04
discuss.huggingface.co
ai-infrastructure
Is there an response length limit for the inference API?
Hugging Face's Inference API returns only 2-3 sentences per response regardless of the model used, according to a user report. The issue may be resolved by setting the `max_new_tokens` parameter higheβ¦