arXiv:2610.00376v1 Announce Type: new Abstract: We evaluate Jev on ten dataset-defined application labels in CESNET-QUICEXT-25 using only the first ten packets' sizes, directions, and inter-packet times. To the best of our knowledge, this is the first empirical study of general-purpose decision models, represented here by Jev, for application classification of network flows. Across 52,000 records from 26 collection weeks following the training period, 40 fixed labeled examples raise Jev's accuracy from 9.80% to 28.42%. Random Forest and Extra Trees trained on 8,000 records achieve 69.95% and 66.80% and outperform Jev in every week. Increasing Jev's context to 150 examples yields 34.50% on the first test week. On a paired 100-record subset, Jev with 40 examples achieves 29% accuracy at a median request time of 0.750 s, versus 37% and 6.036 s for the generative language model OpenAI GPT-5.6 Sol with high reasoning effort through Azure; Jev also incurs lower API charges. The paired subset does not establish an accuracy advantage for either service, and the timing reflects different service configurations. Thus, labeled examples substantially improve Jev, but the tested Jev configurations remain less accurate than trained tree ensembles; unequal supervision budgets and fixed configurations prevent attributing the gap to a single cause.
A First Glance at Jev for Network Traffic Classification: Accuracy, Processing Time, and Cost
A first empirical study of general-purpose decision models for network traffic classification found that Jev reached 28.42% accuracy on ten CESNET-QUICEXT-25 application labels after 40 fixed labeled examples, up from 9.80%, while Random Forest and Extra Trees trained on 8,000 records hit 69.95% and 66.80% and beat Jev in every one of 26 collection weeks. On a paired 100-record subset, Jev with 40 examples scored 29% at a median request time of 0.750 s versus 37% and 6.036 s for OpenAI GPT-5.6 Sol with high reasoning effort through Azure, with Jev incurring lower API charges. The authors report that increasing Jev's context to 150 examples yielded 34.50% on the first test week, and that unequal supervision budgets and fixed configurations prevent attributing the accuracy gap to a single cause.
Run your AI side-project on zahid.host
EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.