Skip to content
Tech News
← Back to articles

Petals: Run LLMs at home, BitTorrent-style

read original more articles
Why This Matters

Petals introduces a decentralized approach to running large language models at home, allowing users to leverage consumer-grade hardware and collaborative networks. This innovation democratizes access to powerful AI models, reducing reliance on centralized APIs and enhancing customization and privacy. It signifies a shift towards more distributed AI infrastructure, empowering individual users and small developers.

Key Takeaways

Petals

Run large language models at home, BitTorrent‑style

Generate text with Llama 3.1 (up to 405B), Mixtral (8x22B), Falcon (40B+) or BLOOM (176B) and fine‑tune them for your tasks — using a consumer-grade GPU or Google Colab.

(up to 405B), (8x22B), (40B+) or (176B) and fine‑tune them for your tasks — using a consumer-grade GPU or Google Colab. You load a part of the model, then join a network of people serving its other parts. Single‑batch inference runs at up to 6 tokens/sec for Llama 2 (70B) and up to 4 tokens/sec for Falcon (180B) — enough for chatbots and interactive apps.

for (70B) and up to for (180B) — enough for chatbots and interactive apps. Beyond classic LLM APIs — you can employ any fine-tuning and sampling methods, execute custom paths through the model, or see its hidden states. You get the comforts of an API with the flexibility of PyTorch and 🤗 Transformers.

Thanks for subscribing! We will email you only if we have really exciting updates.

Top contributors right now:

Loading... Oops, can't load the network status... • and and more

Follow development in Discord or via email: Subscribe We send updates once a few months. No spam. Submitting... We sent you an email to confirm your address. Click it and you're in!

Featured on:

... continue reading