PDF to markdown for LLMs
A developer tested three open-source PDF-to-markdown converters — Microsoft's MarkItDown 0.1.8, IBM Research's Docling 2.131.0, and Datalab's Marker 2.0.0 — on three synthetic PDFs (a two-page report …
A developer tested three open-source PDF-to-markdown converters — Microsoft's MarkItDown 0.1.8, IBM Research's Docling 2.131.0, and Datalab's Marker 2.0.0 — on three synthetic PDFs (a two-page report …
IBM Research and CoreWeave Inc. co-designed identity management and workload isolation controls for agent workloads, extending IBM's internal identity systems into CoreWeave and using CoreWeave Sandbo…
Gartner formalized AI SRE as an analyst category in January 2026 and projects that 85% of enterprises will use AI SRE tooling by 2029, up from less than 5% today, according to its Market Guide for AI …
IBM Research found that a ReAct agent running on GPT-4.1 succeeded on 77.4% of runs on average but passed all five runs on only 53.0% of AppWorld's 168 test_normal tasks, a 24.4-percentage-point consi…
IBM Research demonstrated that membership inference attacks can let an attacker determine whether a specific document is present in a Retrieval-Augmented Generation retrieval database using carefully …
IBM Research introduced consistency guidelines in its altk-evolve toolkit, built on a new Consistency Analyzer, to address a 24.4-point consistency gap in LLM agent reliability. A ReAct agent using GP…
A developer's invoice-triage agent that passed 22 test cases three days in a row later filed the same PDF under two different vendors, illustrating what IBM Research calls the "consistency gap." In a …
NASA's Marshall Space Flight Center Impact AI team and IBM Research released INDUS-SDE, a specialized language model pre-trained on NASA's Science Discovery Engine data that reached 78.1% top-1 masked…
NASA and IBM launched the NASA-IBM Lunar Foundation Model on Thursday, an open-source AI model available on Hugging Face that outperforms widely used methods by up to 23% in identifying the moon's key…
IBM and NASA have developed an open source AI foundation model of the Moon, trained on a lunar dataset aggregating more than 30 spatially-aligned layers from nine instruments across four missions, IBM…
IBM Research and Red Hat used the open-source llm-d framework to deploy GLM-5.2, an approximately 753-billion-parameter mixture-of-experts model, on 544 NVIDIA H100 GPUs, serving up to 3,000 concurren…
IBM Research scientists, including Luis Lastras, Jonathan Lenchner, Barry Trager, Mark Squillante, Chai Wah Wu, Ronald Fagin, and collaborators Wojciech Szpankowski and Alexander Gray, published a pap…
IBM released Granite 4.2 on August 25, 2026, a family of open-weight reasoning models in 3B, 8B, and 30B sizes that enterprises can run on their own hardware under an Apache 2.0 license, with no API m…
IBM Research found that giving AI agents more stored memory often hurts performance, with results varying across eight tested models. The study showed that selectively feeding only necessary informati…
IBM released Granite 4.2, a new family of open-source language models in 3B, 8B, and 30B parameter sizes, purpose-built for enterprise agentic workflows with native reasoning, tool calling, and coding…
IBM Research has introduced BenchDrift, an open-source tool that quantifies how much large language model benchmark scores shift when test prompts are rephrased without changing meaning, revealing tha…
GLM-5.3 tied the leading open weights score with 60 on the Artificial Analysis Intelligence Index, matching Kimi K3, and posted a 246-point jump in agentic Elo from 1524 to 1770, second only to Opus 5…
IBM Research's ALT K-Evolve framework shows that the optimal amount of agentic memory varies by model capability, with strong models like DeepSeek-V3.2 (671B MoE) gaining +9.5 percentage points in tas…
A developer team automated root cause analysis (RCA) for bug tickets using a coding agent that writes structured RCAs at fix time. The real value emerged from aggregating these RCAs, enabling the team…
IBM Research's ALTK-Evolve and ACE both enable LLM agents to learn from their own trajectories, but ALTK-Evolve uses fewer tokens by delivering only a small core of high-support guidelines plus task-s…