cd /news/artificial-intelligence/running-llms-on-a-repurposed-kintex-… · home › topics › artificial-intelligence › article
[ARTICLE · art-148996] src=forum.level1techs.com ↗ pub= topic=artificial-intelligence verified=true sentiment=↑ positive

Running LLMs on a repurposed Kintex-7 PCIe card: openTPU

An open-source project called openTPU is running local LLM inference on a repurposed Inspur YPCB-00338 FPGA accelerator card built around a Kintex-7 xc7k480t with 4 GiB of DDR3 across two channels and PCIe Gen2 x8, hosted in an i7-4790 system. The project reports LFM2.5-230M at 82.1 tokens/s and Qwen3-0.6B at 30.7 tokens/s including host overhead, using 4-bit weights and an int8 LM head, measured over 64 greedy tokens after a 512-token prompt. Decode is largely memory-bandwidth limited, with several larger models reaching 91–94% of the two DDR3 channels' 17.1 GB/s theoretical peak, and the RTL, compiler, simulator, and profiler are released under Apache-2.0, with building requiring a Vivado license covering the device.

read1 min views1 publishedOct 11, 2026
Running LLMs on a repurposed Kintex-7 PCIe card: openTPU
Image: Forum (auto-discovered)

I’ve been turning an Inspur YPCB-00338 FPGA accelerator card into an open-source local LLM inference device.

The card has a Kintex-7 xc7k480t, 4 GiB of DDR3 across two channels, and PCIe Gen2 x8. The current setup runs in an i7-4790 host. The software stack includes a chat client, card monitoring tools, and a profiler; the RTL, compiler, and simulator are all in the same Apache-2.0 repository.

For two concrete physical-card results, LFM2.5-230M runs at 82.1 tokens/s and Qwen3-0.6B at 30.7 tokens/s including host overhead, using 4-bit weights and an int8 LM head. The benchmark is 64 greedy tokens after a 512-token prompt. Larger dense models also run, with their results and quantization tradeoffs listed in the README. Decode is largely limited by memory bandwidth. Several larger models reach 91–94% of the two DDR3 channels’ 17.1 GB/s theoretical peak, based on the card’s counters. There is also an expert-streaming path for MoE models larger than the card’s memory.

The README has a live chat/monitoring demo. You can try the simulator without hardware, and the board guide documents the bitstream, Linux PCIe setup, memory calibration, and diagnostics. Building for this FPGA needs a Vivado license covering the device.

I’d love feedback from people who have worked with surplus PCIe FPGA cards, Linux DMA, or local inference. Board ports and improvements to the bring-up process would be especially useful.

The project is also an experiment in using AI agents for hardware development.

Repo and demo: GitHub - FeSens/openTPU: An open-source AI accelerator, developed by AI: RTL, ISA, simulator, compiler and profiler in one repo. Runs Qwen3, LFM2.5 and Qwen3.5 on a Kintex-7 PCIe card. · GitHub

Board setup: openTPU/docs/board.md at main · FeSens/openTPU · GitHub

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @opentpu 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/running-llms-on-a-re…] indexed:0 read:1min 2026-10-11 · —