Head to head: DeepSeek-V4-Pro vs Phi-4-reasoning DeepSeek-V4-Pro defeated Phi-4-reasoning 12 tasks to 0 with an aggregate score of 105.5 to 33.0 and 100% confidence in a head-to-head benchmark of 12 fresh text tasks scored by gpt-5.4. The evaluation, conducted by an unnamed tester, found DeepSeek-V4-Pro consistently delivered usable outputs in the requested format, while Phi-4-reasoning lost on product discipline, often rambling or ignoring formatting constraints. DeepSeek-V4-Pro didn’t just edge Phi-4-reasoning out — it swept it, 12 tasks to 0 , with an aggregate score of 105.5 to 33.0 and a statistical verdict of 100% confidence . That’s not a narrow technical win or a judge-preference artifact. It’s a decisive result driven by the same pattern over and over: Model A answered the question asked, in the format requested, with clear, usable output. What stands out is how broad the dominance was. DeepSeek-V4-Pro won structured reasoning tasks like budgeting, scheduling, and clinic assignment; coding tasks like the concurrency bug fix and LRU cache; SQL generation; proofreading; contradiction finding; and localization in both Spanish and French Canadian. Even where Model A wasn’t perfect — missing an exact word-count target, slightly awkward phrasing in a translation, or including a bit more explanation than requested — it still produced competent, mostly compliant work. Phi-4-reasoning repeatedly failed at the more basic requirement: actually delivering the final answer cleanly. That’s the story of this matchup. Phi-4-reasoning wasn’t mainly losing on raw intelligence; it was losing on product discipline. In task after task, the judges dinged it for rambling analysis, meta-commentary, exposed reasoning, incomplete outputs, and ignoring explicit formatting constraints like “return only JSON” or “return only the corrected function.” On several problems, it appears to have understood the assignment and even contained the right idea somewhere in the sprawl — but that doesn’t count for much when the user asked for a precise artifact and the model refused to stop talking. For an end user, this distinction is everything. DeepSeek-V4-Pro looks like the model you can drop into real workflows: it gives the SQL, the translation, the schedule, the fix. Phi-4-reasoning looks like a model that too often mistakes process for deliverable. In a benchmark built around practical task completion, that is fatal. Final call: DeepSeek-V4-Pro is the clear winner — not because Phi-4-reasoning had a couple of bad misses, but because it was systematically worse at turning understanding into usable answers. How they were tested We ran 12 fresh text tasks, generated on the fly for this matchup so neither model could prepare in advance, and had gpt-5.4 score each one. To cancel position bias, every task was judged twice — once in each presentation order — and every number reported here, including the headline totals, is the average of both passes. DeepSeek-V4-Pro scored 105.5 to Phi-4-reasoning's 33.0. 1. Find the contradiction The following spec contains exactly one internal contradiction. Quote the two conflicting sentences verbatim and explain the conflict in one sentence. Do not fix it. Spec: "Free accounts may create up to three projects. Every account, regardless of tier, may archive unlimited projects. Archiving a project does not count against the project limit. Free accounts are limited to three projects total, including archived ones." Winner: DeepSeek-V4-Pro — Model A directly quotes the correct two conflicting sentences and gives the required one-sentence explanation. Model B includes extensive unnecessary reasoning, does not cleanly provide the final answer in the requested format, and introduces irrelevant discussion about disclaimers and instructions. Second judge pass, order swapped — scores are the average of both: Model A is better because it directly quotes the two conflicting sentences and gives a one-sentence explanation exactly as requested. Model B identifies the same contradiction correctly, but it adds a large amount of unnecessary reasoning and meta-commentary, so it follows the instructions less precisely and is less cleanly written. 2. support thread 90word summary Summarize the following support thread in exactly 3 bullet points, with exactly 30 words per bullet, faithful to the facts and without adding recommendations. Thread: Mara Ops : Since Tuesday’s 18:40 deploy, barcode scanners in Dock 4 intermittently freeze for 6–8 seconds after a successful scan. Docks 1–3 unaffected. Jin Warehouse lead : Happens on Zebra TC52 units only. We tested 11 devices; 7 reproduced it. Frequency increased during the 07:00–09:00 rush. Leah Backend : Error spikes line up with calls to /inventory/commit. P95 latency jumped from 220 ms to 2.4 s after we enabled duplicate-scan auditing. Rafi Mobile : App logs show UI thread blocking while waiting for commit confirmation. Version 3.18.1 only; 3.17.9 doesn’t freeze. Leah Backend : Temporary mitigation at 10:15 today: disabled duplicate-scan auditing for Dock 4 tenant. Latency dropped to 260 ms; no freezes reported for 95 minutes. Mara Ops : Please note one side effect: audit exports for Dock 4 will be incomplete until the flag is restored. Winner: DeepSeek-V4-Pro — Model A is a concise, faithful summary of the thread, but it fails the exact word-count requirement because its bullets are not 30 words each. Model B does not perform the requested summarization at all and instead exposes reasoning and prompt restatement, badly violating the instructions. Second judge pass, order swapped — scores are the average of both: Model A is better because it actually provides a concise three-bullet summary faithful to the thread’s main points, whereas Model B mostly exposes chain-of-thought and never delivers a valid final summary. However, Model A still fails the exact 30-words-per-bullet requirement and omits several specifics. 3. Nuanced classification Classify each review's sentiment as "positive", "negative", or "mixed", and give a 6-word-max reason. Return ONLY a JSON array of {"text","label","reason"} in input order. Reviews: "Fast shipping but the fabric feels cheap.", "Absolutely love it, wearing it daily ", "It broke after a week. Refund was quick and painless though." Winner: DeepSeek-V4-Pro — Model A fully satisfies the prompt with a valid JSON array in input order, correct sentiment labels, and concise reasons within the six-word limit. Model B includes extensive extraneous analysis instead of returning only the required JSON array, so despite mostly correct classifications, it badly fails instruction adherence. Second judge pass, order swapped — scores are the average of both: Model A fully follows the instruction to return only a JSON array in input order, with correct sentiment labels and concise reasons. Model B includes extensive extraneous analysis and does not return only a JSON array, which is a major instruction-following failure despite mostly correct classifications. 4. Localization with tone Translate this app onboarding line into natural, friendly European Spanish suitable for a mobile toast keep it under 60 characters, no exclamation marks : "You're all set — your first backup starts tonight." Return only the translation, then the character count in parentheses. Winner: DeepSeek-V4-Pro — Model A provides a concise Spanish translation in the requested output format and stays under 60 characters, though "se hará" is slightly less natural than a present-tense onboarding toast. Model B fails the core instruction by returning extensive meta-reasoning instead of only the translation, despite eventually proposing a plausible option. Second judge pass, order swapped — scores are the average of both: Model A is far better because it returns a concise Spanish translation in the requested format and stays under 60 characters. Model B fails the task by outputting extensive meta-reasoning instead of only the translation and count, despite eventually containing a plausible option. 5. project budget reasoning A team is choosing one of three contractors for a 5-week office retrofit. - Contractor Elm charges a fixed setup fee of $3,200 plus $1,850 per week. - Contractor Harbor charges $2,450 per week, with a 12% discount applied to the total weekly charges only if the project lasts at least 5 weeks. - Contractor North charges $1,600 per week plus $2,100 for permits and $900 for cleanup. The company also receives a one-time municipal rebate of $1,500, but only for options whose pre-rebate total exceeds $10,000. Question: After applying any eligible discount and then any eligible rebate, which contractor is cheapest for a 5-week project, and what is the final total cost? Show the calculation clearly. Winner: DeepSeek-V4-Pro — Model A is fully correct, clearly structured, and directly answers the question with concise calculations. Model B eventually reaches the same result, but it includes distracting self-talk, arithmetic confusion, and much weaker presentation despite landing on the correct final answer. Second judge pass, order swapped — scores are the average of both: Model A is clearly better: it gives the correct calculations in a clean, concise format and directly answers the question. Model B eventually reaches the same conclusion, but it includes distracting self-talk, arithmetic confusion, and repetitive filler that significantly hurts clarity and polish. 6. Constraint scheduling Four talks A, B, C, D fill four 1-hour slots 9,10,11,12. Constraints: A is before D; C is not first; B is immediately after A; D is not at 12. Give the ONE valid schedule as 'slot: talk' lines, then a one-line justification. If impossible, say so and explain. Winner: DeepSeek-V4-Pro — Model A and Model B both reach the correct unique schedule, but Model A is better because it is concise and much closer to the requested output format. Model B includes extensive unnecessary internal-style narration and repetition, which weakens instruction adherence and writing quality despite being correct. Second judge pass, order swapped — scores are the average of both: Model A gives the correct unique schedule with a clear, concise justification and clean formatting. Model B is also correct, but it is excessively verbose, includes lots of unnecessary meta-commentary, and does not adhere as tightly to the requested output format. 7. Concurrency bug fix This TypeScript function is meant to memoize an async loader but has a race: concurrent callers can each trigger the underlying fetch. Fix it so the fetch runs at most once per key, and a rejected fetch does NOT poison the cache a later call must retry . Return ONLY the corrected function. ts const cache = new Map