cd /news/generative-ai/report-not-working-3 · home topics generative-ai article
[ARTICLE · art-88759] src=discuss.huggingface.co ↗ pub= topic=generative-ai verified=true sentiment=· neutral

Report: Not working 3

A user attempting to run the tencent/SongGeneration and daydreamlive/DreamVAE models on Hugging Face's free hosted API was told these models cannot be hosted due to their massive memory requirements and specialized inference architecture. The tencent/SongGeneration model has billions of parameters, while DreamVAE requires about 7.5 GB of VRAM and 27 GB of RAM, exceeding free infrastructure limits. The response suggests renting a GPU, using a local PC with a powerful GeForce GPU, or finding a lighter alternative model.

read2 min views1 publishedAug 8, 2026
Report: Not working 3
Image: Discuss (auto-discovered)

I’m sure it wouldn’t be too hard to fix that error, but could you please describe the functionality you want to implement in that space using “natural language” rather than just model names? Just a few lines would be fine… Otherwise, while I can fix the error itself, I won’t be able to make it work properly.

The tencent/SongGeneration model is primarily designed as a large-scale framework (LeLM and music codec) with parameters extending into the billions. Because of its immense memory requirements and specialized inference architecture, it cannot be hosted on Hugging Face’s free CPU-based Inference API.

The daydreamlive/DreamVAE model is not available through the free hosted API because it requires specific, dedicated hardware. Inference requires a robust machine with around ~ 7.5 GB of VRAM and 27 GB of RAM, which exceeds the limitations of the free Hugging Face infrastructure.

Yeah. That’s right. And that’s a problem that can’t be solved even with the best programming or coding. Generative AI is very powerful, but “it’s not magic. It’s technology.” It’s technology that makes it easy to achieve what’s possible. So, it can’t make the impossible possible. By the way, even without generative AI, it’s impossible for me, at least, to make the impossible possible.

So, as a minimum solution, you’ll probably need to rent a GPU from HF for a fee, or if that’s not possible, you’ll need to find some kind of existing API. However, APIs aren’t usually free either.

I tried to download the models from their repos.

Oh, I see. Running OSS models on a local PC is one option. If you have a PC with a really powerful GeForce GPU, you could probably get it to work with enough effort…

But if it’s like my PC’s GPU with only 8GB of VRAM, it might run, but it’ll be super slow…

VRAM capacity isn’t the only indicator of GPU performance, but when it comes to generative AI, running out of VRAM is the biggest hurdle you’ll face. In that case, it’s pretty hopeless—it’s easier to just buy a new GPU or rent one on the cloud.

Or, if a model is too large to run, you could look for a different, lighter model with a similar purpose.

── more in #generative-ai 4 stories · sorted by recency
── more on @hugging face 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/report-not-working-3] indexed:0 read:2min 2026-08-08 ·