{"slug": "running-llms-on-a-repurposed-kintex-7-pcie-card-opentpu", "title": "Running LLMs on a repurposed Kintex-7 PCIe card: openTPU", "summary": "An open-source project called openTPU is running local LLM inference on a repurposed Inspur YPCB-00338 FPGA accelerator card built around a Kintex-7 xc7k480t with 4 GiB of DDR3 across two channels and PCIe Gen2 x8, hosted in an i7-4790 system. The project reports LFM2.5-230M at 82.1 tokens/s and Qwen3-0.6B at 30.7 tokens/s including host overhead, using 4-bit weights and an int8 LM head, measured over 64 greedy tokens after a 512-token prompt. Decode is largely memory-bandwidth limited, with several larger models reaching 91–94% of the two DDR3 channels' 17.1 GB/s theoretical peak, and the RTL, compiler, simulator, and profiler are released under Apache-2.0, with building requiring a Vivado license covering the device.", "body_md": "I’ve been turning an Inspur YPCB-00338 FPGA accelerator card into an open-source local LLM inference device.\n\nThe card has a Kintex-7 xc7k480t, 4 GiB of DDR3 across two channels, and PCIe Gen2 x8. The current setup runs in an i7-4790 host. The software stack includes a chat client, card monitoring tools, and a profiler; the RTL, compiler, and simulator are all in the same Apache-2.0 repository.\n\nFor two concrete physical-card results, LFM2.5-230M runs at 82.1 tokens/s and Qwen3-0.6B at 30.7 tokens/s including host overhead, using 4-bit weights and an int8 LM head. The benchmark is 64 greedy tokens after a 512-token prompt. Larger dense models also run, with their results and quantization tradeoffs listed in the README.\n\nDecode is largely limited by memory bandwidth. Several larger models reach 91–94% of the two DDR3 channels’ 17.1 GB/s theoretical peak, based on the card’s counters. There is also an expert-streaming path for MoE models larger than the card’s memory.\n\nThe README has a live chat/monitoring demo. You can try the simulator without hardware, and the board guide documents the bitstream, Linux PCIe setup, memory calibration, and diagnostics. Building for this FPGA needs a Vivado license covering the device.\n\nI’d love feedback from people who have worked with surplus PCIe FPGA cards, Linux DMA, or local inference. Board ports and improvements to the bring-up process would be especially useful.\n\nThe project is also an experiment in using AI agents for hardware development.\n\nRepo and demo: [GitHub - FeSens/openTPU: An open-source AI accelerator, developed by AI: RTL, ISA, simulator, compiler and profiler in one repo. Runs Qwen3, LFM2.5 and Qwen3.5 on a Kintex-7 PCIe card. · GitHub](https://github.com/FeSens/openTPU)\n\nBoard setup: [openTPU/docs/board.md at main · FeSens/openTPU · GitHub](https://github.com/FeSens/openTPU/blob/main/docs/board.md)", "url": "https://wpnews.pro/news/running-llms-on-a-repurposed-kintex-7-pcie-card-opentpu", "canonical_source": "https://forum.level1techs.com/t/running-llms-on-a-repurposed-kintex-7-pcie-card-opentpu/258369#post_2", "published_at": "2026-10-11 03:39:49+00:00", "updated_at": "2026-10-11 03:50:32.631027+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-infrastructure", "ai-tools", "developer-tools"], "entities": ["openTPU", "Kintex-7 xc7k480t", "Inspur YPCB-00338", "LFM2.5-230M", "Qwen3-0.6B", "Vivado", "Apache-2.0", "GitHub"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/running-llms-on-a-repurposed-kintex-7-pcie-card-opentpu", "markdown": "https://wpnews.pro/news/running-llms-on-a-repurposed-kintex-7-pcie-card-opentpu.md", "text": "https://wpnews.pro/news/running-llms-on-a-repurposed-kintex-7-pcie-card-opentpu.txt", "jsonld": "https://wpnews.pro/news/running-llms-on-a-repurposed-kintex-7-pcie-card-opentpu.jsonld"}}