Head to head: DeepSeek-V3.2 vs Phi-4 DeepSeek-V3.2 defeated Phi-4 in a head-to-head benchmark with an aggregate score of 98.0 to 81.8, an 8–2 task lead with 2 ties, and a 97% confidence verdict, according to an evaluation by gpt-5.4 across 12 fresh text tasks. DeepSeek-V3.2 won on reasoning accuracy, arithmetic, and instruction-following, while Phi-4's wins were narrower and both models missed the O(1) requirement in the TypeScript LRU task. DeepSeek-V3.2 takes this head-to-head decisively: a 98.0 to 81.8 aggregate score, an 8–2 task lead with 2 ties , and a 97% confidence verdict. That is not a vibes-based edge or a split-decision win; it’s a broad, statistically convincing result driven by stronger reasoning accuracy and fewer costly mistakes. What stands out is where DeepSeek-V3.2 won. It was better on contradiction finding, unit-aware math, the clinic rota puzzle, proofreading, quarterly budget calculation, tightly formatted status writing, and nuanced classification. In several of those, the difference wasn’t style — it was correctness. Phi-4 lost ground by muddying the core contradiction, inventing uniqueness in a non-unique rota, making an arithmetic error in the budget check, and overrunning explicit constraints in the classification task. Phi-4 did earn real wins, but they were narrower. It was better at messy vendor email parsing because it preserved more structure and source detail, and it edged the TypeScript LRU task mostly on cleaner typing — though both models missed the stated O 1 requirement there, which limits how much credit either deserves. The ties tell a similar story: on strict JSON extraction, each model dropped a different part of the brief, and on the LATAM Spanish SMS, both had tradeoffs between precision, idiom, and length control. The bigger editorial takeaway is that DeepSeek-V3.2 was the sturdier model under pressure. Even when it wasn’t perfect — as in the rota task, where it correctly spotted multiple valid solutions but then overcommitted to one — it still showed better judgment about the underlying problem. Phi-4 was more likely to look competent while slipping on the thing that actually mattered, whether that was arithmetic, constraint handling, or identifying the precise source of a contradiction. Final call: DeepSeek-V3.2 is the clear winner. Not because it dazzled in one or two showcase prompts, but because it was more dependable across reasoning, calculation, editing, and instruction-following — and the numbers back that up emphatically. How they were tested We ran 12 fresh text tasks, generated on the fly for this matchup so neither model could prepare in advance, and had gpt-5.4 score each one. To cancel position bias, every task was judged twice — once in each presentation order — and every number reported here, including the headline totals, is the average of both passes. DeepSeek-V3.2 scored 98.0 to Phi-4's 81.8. 1. Find the contradiction The following spec contains exactly one internal contradiction. Quote the two conflicting sentences verbatim and explain the conflict in one sentence. Do not fix it. Spec: "Free accounts may create up to three projects. Every account, regardless of tier, may archive unlimited projects. Archiving a project does not count against the project limit. Free accounts are limited to three projects total, including archived ones." Winner: DeepSeek-V3.2 — Model A better identifies the contradiction by pairing the two conflicting sentences and explaining the clash without introducing extra ambiguity. Model B quotes the same pair, but its explanation leans on the separate archived-project sentence rather than cleanly explaining the direct conflict between the quoted lines. Second judge pass, order swapped — scores are the average of both: Model A better identifies and explains the actual contradiction by tying the first sentence to the implication that archived projects could be replaced with new ones, which directly conflicts with the total-including-archived cap. Model B quotes the same pair, but its explanation is less precise and muddles the conflict by bringing in an additional sentence without clearly centering the contradiction. 2. Strict JSON extraction Extract every meeting from this text into a JSON array of {"title","day","start24","durationMinutes"}. Use 24-hour "HH:MM" for start24, integer minutes for duration, and title-case titles. Return ONLY valid JSON, no commentary. Text: "Standup is daily at 9am for a quarter hour. The design review runs thursday from 2:30-4pm. Payroll sync — first monday, 11:00 to noon." Winner: Tie — Model A is better because it returns valid JSON only and correctly extracts the times and durations. Model B has a more accurate day value for "First Monday," but it violates the instruction to return only valid JSON by wrapping the output in a code fence. Second judge pass, order swapped — scores are the average of both: Model B is better because it correctly preserves the recurrence/day information as title-cased values like "Daily," "Thursday," and "First Monday," while Model A lowercases day values and loses the "First" qualifier for the payroll meeting. Both parse times and durations correctly, but Model B adheres more closely to the formatting and extraction requirements. 3. Unit-aware math A pump moves 3.5 liters every 8 seconds. A tank holds 0.9 cubic meters. Starting empty, how long to fill it, in minutes and seconds mm:ss , rounded to the nearest second? Show the key steps, then give the final answer on its own line. Winner: DeepSeek-V3.2 — Model A and Model B both reach the correct final answer, but Model A is slightly better because it explicitly checks the rounding to the nearest second and presents the key steps more rigorously. Model B is also strong, but its final answer is not placed quite as cleanly on its own line after the explanation. Second judge pass, order swapped — scores are the average of both: Model A is slightly better because it is equally correct but handles the rounding step more explicitly and cleanly, which better matches the request for key steps. Model B is also strong, but its conversion from total seconds to mm:ss is a bit less precise in presentation than Model A’s exact-seconds check before rounding. 4. Clinic rota assignment Assign four staff members to four clinic tasks for Tuesday morning, using each person exactly once. People: Nia, Omar, Priya, Luis Tasks: Reception, Triage, Vaccines, Inventory Constraints: - Nia cannot do Reception. - Omar must do either Triage or Inventory. - Priya cannot do Vaccines. - Luis must do Reception or Vaccines. - If Omar does Inventory, then Priya must do Reception. - Reception cannot be assigned to Nia or Omar. - Triage must be assigned to either Omar or Priya. Find the unique valid assignment and present it as Person - Task on separate lines. Winner: DeepSeek-V3.2 — Model A correctly notices that the stated constraints actually permit two valid assignments, so the puzzle as written is not uniquely solvable; however, it still undermines itself by guessing one final assignment instead of cleanly reporting the inconsistency. Model B makes an invalid inference that Luis must be Reception and presents one of the two valid assignments as unique, so its reasoning is less correct overall. Second judge pass, order swapped — scores are the average of both: Model A is better because it correctly identifies that the stated constraints actually permit two valid assignments and explains this carefully, whereas Model B incorrectly claims uniqueness and gives just one of the valid solutions. Model B follows the requested output format better, but Model A is substantially stronger on correctness and reasoning. 5. Messy vendor email parsing Extract structured facts from this messy email thread. Return a JSON object with exactly these keys: vendor, event date, venue, arrival time, headcount, dietary counts, total quote usd, deposit due usd, contact email, contact phone. Email text: "From: Miri Calder miri@harborandpine.co mailto:miri@harborandpine.co Subject: Re: June workshop catering Hi Jae — confirming Harbor & Pine can cater the product workshop on Fri, Sept 12 at Studio North, 18 Vale St. We should arrive by 8:15 a.m. for a 9:00 start. Current count is 27 attendees + 3 facilitators = 30 total. Food notes from your form: 4 vegetarian, 2 vegan, 1 gluten-free. No shellfish. Quote v2 attached in my last note: breakfast package $540, coffee service $135, staffing $90, delivery waived. Total = USD 765. To lock it in, we need a 40% deposit by Aug 29. Best contact day-of is me at 555-0188, or reply here. —Miri p.s. if AP needs it, our billing alias is invoices@harborandpine.co mailto:invoices@harborandpine.co " Winner: Phi-4 — Model B is better because it preserves more source detail in the date and venue and structures dietary information more usefully while remaining accurate. Model A is mostly correct, but its dietary counts is an unstructured string and it omits minor punctuation/detail from the source. Second judge pass, order swapped — scores are the average of both: Model B is better because it preserves more structure from the email and represents dietary information in a more usefully parsed form while keeping all required keys. Model A is mostly correct, but its dietary counts value is an unstructured string rather than extracted structured facts, making it less aligned with the task. 6. LRU cache Implement a class LRUCache