I built a 17-agent AI swarm on my phone — here's how A developer built a 17-agent AI swarm that runs entirely on a Pixel 7 smartphone, using llama-cpp-python and a quantized Phi-2 model to analyze long-form text and generate tailored social media posts for X, LinkedIn, Facebook, and Instagram. The system splits the task across specialized agents for summarization, keyword extraction, and per-platform post generation, with GPU layers offloaded to the phone's processor. Okay, buckle up. This is going to be a bit of a deep dive. For the past few weeks I've been obsessing over the idea of running a genuinely useful, multi-agent system entirely on a smartphone. Not just a demo, not just a simplified example, but something that could actually perform a relatively complex task. And I did it. I built a 17-agent AI swarm, running on my Pixel 7, that analyzes text and generates targeted social media content. Why a phone? Honestly, it's a challenge. We're so used to thinking of AI as cloud-based, reliant on massive servers. But the capabilities of modern smartphone processors are frankly astonishing . It forces you to be incredibly efficient, to think about model size, quantization, and creative code architecture. Plus, it’s portable. Where else can you carry a swarm intelligence around in your pocket? This isn't about replacing large language models LLMs hosted in the cloud. It's about exploring the boundaries of what’s possible on-device . The Core Idea: Social Media Content Alchemy The goal was to take a relatively long-form piece of text think a blog post, article, or even a transcript and automatically generate a set of tailored social media posts for different platforms – Twitter now X , LinkedIn, Facebook, and Instagram. The 'swarm' architecture was crucial. Instead of relying on one big model, I broke the problem down into specialized agents, each handling a specific stage of the process. Think of it like an assembly line, but powered by AI. The Agents: A Breakdown of the Swarm Here’s a look at the 17 agents and their roles: Tech Stack & Challenges transformers and llama-cpp-python libraries. Code Snippets Illustrative Let's look at some simplified snippets to give you a flavour. 1. Loading the Model Python - llama-cpp-python python from llama cpp import Llama llm = Llama model path="./phi-2.Q4 K M.gguf", n ctx=2048, n gpu layers=-1 Use all available GPU layers This initializes the Llama model, loading the quantized GGUF file. The -1 argument tries to offload as many layers as possible to the GPU which my Pixel 7 has . 2. A Simple Agent Function Post Generator python def generate post summary, keywords, platform, length : prompt = f"Generate a {length} social media post for {platform} about the following:\n\nSummary: {summary}\n\nKeywords: {keywords}\n\nPost:" output = llm prompt, max tokens=150, stop= "\n\n" , echo=False return output 'choices' 0 'text' .strip This function takes the summarized text, keywords, platform, and desired post length as input and constructs a prompt for Phi-2. The stop parameter prevents the model from generating endlessly. 3. Orchestrating the Swarm Simplified python Load input text input text = load text from file "my article.txt" Summarize using the Summarizer agents summaries = summarize text input text for in range 2 Extract Keywords using the Keyword Extractor agents keywords sets = extract keywords input text for in range 2 Average the keyword sets keywords = list set keywords sets 0 + keywords sets 1 Remove Duplicates Generate posts for each platform platform posts = {} for platform in "Twitter", "LinkedIn", "Facebook", "Instagram" : platform posts platform = generate post summaries 0 , keywords, platform, "short" , generate post summaries 0 , keywords, platform, "medium" , generate post summaries 0 , keywords, platform, "long" Print Results for platform, posts in platform posts.items : print f"--- {platform} ---" for i, post in enumerate posts : print