The llama.cpp and vLLM projects move INSANELY fast. do we have a good grasp on how good tensor-parallel inference on Intel Arc is at the end of Aug 2026? As of late August 2026, the llama.cpp and vLLM projects are advancing rapidly, but community members report that the most recent information on tensor-parallel inference for Intel Arc GPUs remains a Reddit post from earlier in the year, which highlighted a merged PR for TP SYCL in llama.cpp. Users are seeking updated performance benchmarks, particularly for running Qwen3.6/8-27B on dual Arc Pro B65 GPUs, with some considering switching from two 16GB 5060Tis to gain sufficient VRAM for reasonable quantization without quantizing KV cache. this is the most recent stuff I can find: Reddit https://www.reddit.com/r/LocalLLaMA/comments/1ufctc8/tensor split fix for intel gpus llamacpp release/ everything more recent I find just ends up linking back to that post. which like, sure, I get it, that was where folks pointed out the PR for TP SYCL that got merged into llama.cpp. I just want to know if folks around here are running it and whether or not you’ve seen improvements in performance over the past two months, especially on everyone’s favorite, Qwen3.6/8-27B. I am slowly leaning in the direction of selling my two 16GB 5060Tis and purchasing a couple of Arc Pro B65s if I can find them in order to have the VRAM to run that model specifically at a reasonable quant without quantizing kv.