I’ve been turning an Inspur YPCB-00338 FPGA accelerator card into an open-source local LLM inference device.
The card has a Kintex-7 xc7k480t, 4 GiB of DDR3 across two channels, and PCIe Gen2 x8. The current setup runs in an i7-4790 host. The software stack includes a chat client, card monitoring tools, and a profiler; the RTL, compiler, and simulator are all in the same Apache-2.0 repository.
For two concrete physical-card results, LFM2.5-230M runs at 82.1 tokens/s and Qwen3-0.6B at 30.7 tokens/s including host overhead, using 4-bit weights and an int8 LM head. The benchmark is 64 greedy tokens after a 512-token prompt. Larger dense models also run, with their results and quantization tradeoffs listed in the README. Decode is largely limited by memory bandwidth. Several larger models reach 91–94% of the two DDR3 channels’ 17.1 GB/s theoretical peak, based on the card’s counters. There is also an expert-streaming path for MoE models larger than the card’s memory.
The README has a live chat/monitoring demo. You can try the simulator without hardware, and the board guide documents the bitstream, Linux PCIe setup, memory calibration, and diagnostics. Building for this FPGA needs a Vivado license covering the device.
I’d love feedback from people who have worked with surplus PCIe FPGA cards, Linux DMA, or local inference. Board ports and improvements to the bring-up process would be especially useful.
The project is also an experiment in using AI agents for hardware development.
Board setup: openTPU/docs/board.md at main · FeSens/openTPU · GitHub