# I built a 17-agent AI swarm on my phone — here's how

> Source: <https://dev.to/sam_hiotis_117598dbfa3ac2/i-built-a-17-agent-ai-swarm-on-my-phone-heres-how-fb8>
> Published: 2026-10-11 08:39:31+00:00

Okay, buckle up. This is going to be a bit of a deep dive. For the past few weeks I've been obsessing over the idea of running a genuinely useful, multi-agent system *entirely* on a smartphone. Not just a demo, not just a simplified example, but something that could actually perform a relatively complex task. And I did it. I built a 17-agent AI swarm, running on my Pixel 7, that analyzes text and generates targeted social media content. 

Why a phone? Honestly, it's a challenge. We're so used to thinking of AI as cloud-based, reliant on massive servers. But the capabilities of modern smartphone processors are frankly *astonishing*. It forces you to be incredibly efficient, to think about model size, quantization, and creative code architecture. Plus, it’s portable. Where else can you carry a swarm intelligence around in your pocket?

This isn't about replacing large language models (LLMs) hosted in the cloud. It's about exploring the boundaries of what’s possible *on-device*.

**The Core Idea: Social Media Content Alchemy**

The goal was to take a relatively long-form piece of text (think a blog post, article, or even a transcript) and automatically generate a set of tailored social media posts for different platforms – Twitter (now X), LinkedIn, Facebook, and Instagram.

The 'swarm' architecture was crucial. Instead of relying on one big model, I broke the problem down into specialized agents, each handling a specific stage of the process. Think of it like an assembly line, but powered by AI.

**The Agents: A Breakdown of the Swarm**

Here’s a look at the 17 agents and their roles:

**Tech Stack & Challenges**

`transformers` and `llama-cpp-python` libraries.
**Code Snippets (Illustrative)**

Let's look at some simplified snippets to give you a flavour.

**1. Loading the Model (Python - `llama-cpp-python`)**

``` python
from llama_cpp import Llama

llm = Llama(model_path="./phi-2.Q4_K_M.gguf", n_ctx=2048, n_gpu_layers=-1)  # Use all available GPU layers
```

This initializes the Llama model, loading the quantized GGUF file. The `-1` argument tries to offload as many layers as possible to the GPU (which my Pixel 7 has).

**2.  A Simple Agent Function (Post Generator)**

``` python
def generate_post(summary, keywords, platform, length):
    prompt = f"Generate a {length} social media post for {platform} about the following:\n\nSummary: {summary}\n\nKeywords: {keywords}\n\nPost:"
    output = llm(prompt, max_tokens=150, stop=["\n\n"], echo=False)
    return output['choices'][0]['text'].strip()
```

This function takes the summarized text, keywords, platform, and desired post length as input and constructs a prompt for Phi-2.  The `stop` parameter prevents the model from generating endlessly.  

**3. Orchestrating the Swarm (Simplified)**

```
python
# Load input text
input_text = load_text_from_file("my_article.txt")

# Summarize (using the Summarizer agents)
summaries = [summarize_text(input_text) for _ in range(2)]

# Extract Keywords (using the Keyword Extractor agents)
keywords_sets = [extract_keywords(input_text) for _ in range(2)]

# Average the keyword sets
keywords = list(set(keywords_sets[0] + keywords_sets[1])) #Remove Duplicates

# Generate posts for each platform
platform_posts = {}
for platform in ["Twitter", "LinkedIn", "Facebook", "Instagram"]:
    platform_posts[platform] = [
        generate_post(summaries[0], keywords, platform, "short"),
        generate_post(summaries[0], keywords, platform, "medium"),
        generate_post(summaries[0], keywords, platform, "long")
    ]

# Print Results
for platform, posts in platform_posts.items():
    print(f"--- {platform} ---")
    for i, post in enumerate(posts):
        print(
```


