cd /news/large-language-models/splitting-documents-at-lower-cost-mu… · home topics large-language-models article
[ARTICLE · art-137822] src=arxiv.org ↗ pub= topic=large-language-models verified=true sentiment=· neutral

Splitting Documents at Lower Cost: Multi-Split Boundary Decisions for LLM-Based Page Stream Segmentation

A new arXiv paper (2609.22620v1) introduces Multi-Split Boundary Decision (MSBD), a method that predicts multiple document boundaries within a page window in a single large language model call, reducing inference requests compared with standard Page Classification and Boundary Decision formulations that resolve only one boundary per call. Evaluating MSBD across multiple language models, document collections, input modalities, and window sizes, the authors found a model- and corpus-dependent operating range where MSBD preserves strong segmentation accuracy while substantially improving inference efficiency, followed by a sharp decline at larger windows. MSBD delivered the strongest overall accuracy-efficiency trade-off, with large windows exposing distinct over- and under-segmentation behavior across models.

by read1 min views1 publishedSep 23, 2026

arXiv:2609.22620v1 Announce Type: new Abstract: Scanned mail, uploaded PDFs, and consolidated attachments often arrive as page streams that must be split into individual documents before downstream classification, extraction, or routing. Zero-shot large language models can detect document boundaries without task-specific training, but standard Page Classification (PC) and Boundary Decision (BD) formulations resolve only one boundary per model call. We introduce Multi-Split Boundary Decision (MSBD), which predicts multiple boundaries within a page window in a single call, reducing the number of inference requests. We evaluate MSBD across multiple language models, document collections, input modalities, and window sizes. The results reveal a model- and corpus-dependent operating range in which MSBD preserves strong segmentation accuracy while substantially improving inference efficiency, followed by a sharp decline at larger windows. MSBD provided the strongest overall accuracy--efficiency trade-off, while large windows expose distinct over- and under-segmentation behavior across models. These findings show that multi-boundary prediction can make zero-shot page stream segmentation more efficient when the window size is selected for the target corpus.

── more in #large-language-models 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/splitting-documents-…] indexed:0 read:1min 2026-09-23 ·