cd /news/machine-learning/kuafu-compressing-long-user-behavior… · home › topics › machine-learning › article
[ARTICLE · art-140843] src=machinebrief.com ↗ pub= topic=machine-learning verified=true sentiment=↑ positive

KuaFu: Compressing Long User Behavior into Understanding at Billion Scale

Tencent researchers presented KuaFu, a unified behavior-compression layer that compresses each user behavior item into 2-4 tokens of width 128-256, cutting per-item cache from 10 KB to 0.5 KB, according to the arXiv paper 2609.31045v1. Across four production profiling tasks KuaFu matched or exceeded uncompressed single-task production models on all five headline metrics, raised per-GPU throughput by 37%-350%, and saved 190 GPUs, and the system has run on Tencent's advertising and recommendation platform for ten months, lifting overall GMV by 1.37%. On public benchmarks KuaFu nearly always beat prior compressors at the same compression ratio, including up to +17.7 EM on out-of-domain MRQA, and on RecBench a 4B model surpassed its 8B counterpart by 1.90 points.

by read1 min views1 publishedSep 28, 2026

arXiv:2609.31045v1 Announce Type: cross Abstract: Conversational agents, generative recommenders, and personalized advertising all rest on one capability: understanding each user from raw behavior. Prevailing industrial practice is task-specific: for each task, a relevant subsequence is extracted from the full history and a dedicated model trained on it. In production it hits two bottlenecks. First, even after filtering, a single-task sequence stays extremely long: content-interest summarization reads several hundred items per user, tens of thousands of tokens once serialized as prompt text. Second, profiles are refreshed routinely: a billion users weekly, roughly 100K QPM in aggregate, which under a fixed GPU budget sets a hard throughput floor. Compression is therefore mandatory, yet truncation or coarse compression can silently distort the profile, introducing four hallucination types (fabrication, omission, date misattribution, broken logic) that, with no way to evaluate the compressed representation itself, surface only as diffuse degradation in downstream metrics. We present KuaFu, a unified behavior-compression layer whose minimal unit is one behavior item. A two-axis projector compresses each item into 2-4 tokens of width 128-256 (about 10x along the token axis, 20x along width; per-item cache 10 KB to 0.5 KB), with fidelity-oriented four-stage training and layered intermediate evaluation. Across four production profiling tasks it matches or exceeds uncompressed single-task production models on all five headline metrics, raises per-GPU throughput by 37%-350%, and saves 190 GPUs. On public benchmarks it nearly always beats prior compressors at the same compression ratio (up to +17.7 EM on out-of-domain MRQA); on RecBench, a 4B model surpasses its 8B counterpart by 1.90 points. KuaFu has run on the Tencent advertising and recommendation platform for ten months, lifting overall GMV by 1.37%.

── more in #machine-learning 4 stories · sorted by recency
── more on @kuafu 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/kuafu-compressing-lo…] indexed:0 read:1min 2026-09-28 · —