# Run large language models at home, BitTorrent‑style

> Source: <https://petals.dev/>
> Published: 2026-07-23 01:33:12+00:00

# Petals

Run large language models at home, BitTorrent‑style

-
Generate text with
**Llama 3.1**(up to 405B),** Mixtral**(8x22B),** Falcon**(40B+) or** BLOOM**(176B) and fine‑tune them for your tasks — using a consumer-grade GPU or Google Colab. -
You load a part of the model, then join a
[network](https://health.petals.dev)of people serving its other parts. Single‑batch inference runs at up to**6 tokens/sec** for**Llama 2**(70B) and up to** 4 tokens/sec**for** Falcon**(180B) — enough for[chatbots](https://chat.petals.dev)and interactive apps. -
Beyond classic LLM APIs —
you can employ any fine-tuning and sampling methods, execute custom paths through the model, or see its hidden states.
You get the comforts of an API with the flexibility of
**PyTorch** and 🤗**Transformers**.

**Top contributors** right now:

Loading...

Follow development in [Discord](https://discord.gg/D9MwApKgWa) or via email:

We send updates once a few months. No spam.

We sent you an email to confirm your address. Click it and you're in!

Featured on:

This project is a part of the [BigScience](https://bigscience.huggingface.co/) research workshop.
