{"slug": "hss-synth-humanities-and-social-sciences-data-synthesis-for-llms", "title": "HSS-Synth: Humanities and Social Sciences Data Synthesis for LLMs", "summary": "Researchers introduced HSS-Synth, the first data synthesis pipeline for humanities and social sciences, generating 237,000 instruction-tuning samples that outperform 14 leading baselines on 16 benchmarks. The fine-tuned Qwen3-8B-Base model achieved state-of-the-art results, approaching the official Qwen3-8B, with code publicly available on GitHub.", "body_md": "arXiv:2607.27379v1 Announce Type: new\nAbstract: High-quality, diverse data are vital for large language models (LLMs) but remain scarce and costly. Data synthesis is a viable alternative and succeeds on closed tasks, yet the humanities and social sciences (HSS) are overlooked, and their open-ended nature makes synthesis challenging. Moving beyond prior capability-centric, fragmented attempts, we adopt a subject-centric paradigm, define the first HSS domain system covering 14 mainstream fields, and introduce HSS-Synth, the first data synthesis pipeline for HSS. HSS-Synth comprises: (1) constructing seed documents from web corpora via multi-step filtering and text refinement evaluated by a judge; (2) specifying \"requirements + persona\" to backtranslate seed documents into diverse yet faithful instructions with a strict Q&A alignment check; and (3) breaking LLM response limits via teacher-forced Answering that feeds seed documents during response generation to anchor semantics, reduce hallucinations, and preserve tone and integrity. HSS-Synth yields 237k high-quality, diverse instruction-tuning samples that outperform 14 leading baselines on 16 benchmarks. The fine-tuned Qwen3-8B-Base sets a new SOTA and approaches the official Qwen3-8B, improving both human preference and knowledge capabilities without performance seesaws. Extensive experiments demonstrate HSS-Synth's robustness and transferability. Our code is publicly available at https://github.com/pengr/HSS-Synth.", "url": "https://wpnews.pro/news/hss-synth-humanities-and-social-sciences-data-synthesis-for-llms", "canonical_source": "https://www.machinebrief.com/news/hss-synth-humanities-and-social-sciences-data-synthesis-for-0jl1", "published_at": "2026-07-31 04:00:00+00:00", "updated_at": "2026-07-31 04:36:20.387840+00:00", "lang": "en", "topics": ["large-language-models", "artificial-intelligence", "ai-research", "ai-products"], "entities": ["HSS-Synth", "Qwen3-8B-Base", "Qwen3-8B", "arXiv", "GitHub"], "alternates": {"html": "https://wpnews.pro/news/hss-synth-humanities-and-social-sciences-data-synthesis-for-llms", "markdown": "https://wpnews.pro/news/hss-synth-humanities-and-social-sciences-data-synthesis-for-llms.md", "text": "https://wpnews.pro/news/hss-synth-humanities-and-social-sciences-data-synthesis-for-llms.txt", "jsonld": "https://wpnews.pro/news/hss-synth-humanities-and-social-sciences-data-synthesis-for-llms.jsonld"}}