cd /news/large-language-models/show-hn-jeva-cpp-a-llama-cpp-fork-wi… · home › topics › large-language-models › article
[ARTICLE · art-141039] src=github.com ↗ pub= topic=large-language-models verified=true sentiment=· neutral

Show HN: Jeva.cpp – a llama.cpp fork with JEV-compatible API for all LLMs

Jeva.cpp, a fork of llama.cpp, adds a JEV-compatible decision API to llama-server that enables Choice, Score and Noul evaluations directly from model logits while preserving standard autoregressive generation. The fork works with all models and platforms supported by llama.cpp, reusing its existing model implementations and inference backends, though the JEV decision API requires models that provide next-token vocabulary logits.

read3 min views1 publishedSep 28, 2026
Show HN: Jeva.cpp – a llama.cpp fork with JEV-compatible API for all LLMs
Image: Michielbdejong (auto-discovered)

jeva.cpp is a fork of llama.cpp that adds a JEV-compatible decision API to llama-server, enabling Choice, Score and Noul evaluations directly from model logits while preserving standard autoregressive generation.

jeva.cpp is designed to work with all models and platforms supported by llama.cpp, reusing its existing model implementations and inference backends. The JEV decision API requires models that provide next-token vocabulary logits; other model types retain their original functionality.

LLM inference in C/C++

ggml / ops / maintainer PRs / dev stats / lib llama API / llama-server REST API

A few options to get llama.cpp installed on your machine:

Once installed:

llama cli -hf ggml-org/Qwen3.5-0.8B-GGUF

llama serve -hf ggml-org/Qwen3.5-0.8B-GGUF

| VLM session with llama cli | Built-in web UI against llama serve |

The main goal of llama.cpp is to enable LLM (and VLM) inference with minimal setup and state-of-the-art performance on a wide range of hardware - locally and in the cloud.

  • Plain C/C++ implementation without any dependencies
  • Apple silicon is a first-class citizen - optimized via ARM NEON, Accelerate and Metal frameworks
  • AVX, AVX2, AVX512 and AMX support for x86 architectures
  • RVV, ZVFH, ZFH, ZICBOP and ZIHINT support for RISC-V architectures
  • 1.5-bit, 2-bit, 3-bit, 4-bit, 5-bit, 6-bit, and 8-bit integer quantization for faster inference and reduced memory use
  • Custom CUDA kernels for running LLMs on NVIDIA GPUs (support for AMD GPUs via HIP and Moore Threads GPUs via MUSA)
  • Vulkan and SYCL backend support
  • CPU+GPU hybrid inference to partially accelerate models larger than the total VRAM capacity

The llama.cpp project is build on top of the ggml library.

Backend Target devices
BLAS All
BLIS All
CANN Ascend NPU
CUDA Nvidia GPU
HIP AMD GPU
Hexagon Snapdragon
IBM zDNN IBM Z & LinuxONE
MUSA Moore Threads GPU
Metal Apple Silicon
OpenCL Adreno GPU
OpenVINO [In Progress] Intel CPUs, GPUs, and NPUs
RPC All
SYCL Intel GPU
VirtGPU VirtGPU APIR
Vulkan GPU
WebGPU All
ZenDNN AMD CPU
  • Contributors can open PRs

  • Collaborators will be invited based on contributions

  • Maintainers can push to branches in the llama.cpp repo and merge PRs into themaster branch

  • Any help with managing issues, PRs and projects is very appreciated!

  • Read the CONTRIBUTING.md for more information

  • yhirose/cpp-httplib - Single-header HTTP server, used byllama-server - MIT license

  • nothings/stb - Single-header image format decoder, used by multimodal subsystem - Public domain

  • nlohmann/json - Single-header JSON library, used by various tools/examples - MIT License

  • mackron/miniaudio - Single-header audio format decoder, used by multimodal subsystem - Public domain

  • sheredom/subprocess.h - Single-header process launching solution for C and C++ - Public domain

── more in #large-language-models 4 stories · sorted by recency
── more on @jeva.cpp 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/show-hn-jeva-cpp-a-l…] indexed:0 read:3min 2026-09-28 · —