{"slug": "0-mtech-phd-md", "title": "[0] Mtech_phd.md", "summary": "A developer known as @danisherror has published a self-directed M.Tech program in Networked Systems & Security for AI Infrastructure, a two-year, four-semester self-study curriculum of roughly 15 hours per week (~1,500 hours total) aimed at building and securing the infrastructure AI runs on, including GPU datacenter networks, sandboxed agent execution, and AI-assisted network automation. The plan sequences six core courses across three semesters plus three electives, running a heavy main track (operating systems, then networks) alongside a lighter side track (Go, then math), with xv6 labs, MIT 6.5840 material, and datacenter papers such as Al-Fares fat-tree, VL2, and Jupiter Rising. It also specifies a weekly schedule, rest weeks every six weeks, and a fallback minimum of about four hours per week when work gets busy.", "body_md": "Got it, this is your self-directed M.Tech in Networked Systems & Security for AI Infrastructure. That changes things from my earlier generic answer, so here's a schedule built specifically around this program's Semester 1 (C1 OS, C2 Advanced Networks, C3 Math, C4 Go), fitted to a 5-hour workday.\n\nHow much time to use\n\nYour plan is designed for about 15 hours a week. With a 5-hour job you have room for more, but I'd keep 15 hours as the committed load and treat the extra capacity as a buffer (Friday catch-up, longer lab sessions when something breaks). Two years is long, and the plan's own rule of a week off every 6 weeks only works if the weekly load stays sustainable.\n\nThe weekly routine (same slots your doc uses; shift them if your work hours differ)\n\nDay\n\nTime\n\nHrs\n\nWhat\n\nMon\n\n9–11 PM\n\n2\n\nMain course: lectures and reading\n\nTue\n\n9–11 PM\n\n2\n\nMain course: lab coding\n\nWed\n\n9–10 PM\n\n1\n\nPaper of the week (first and second pass)\n\nThu\n\n9–11 PM\n\n2\n\nSide course: lectures, exercises\n\nFri\n\n9–11 PM\n\n0–2\n\nBuffer only: catch up if behind, otherwise rest\n\nSat\n\n11 AM–2 PM\n\n3\n\nDeep work: main course lab or project\n\nSun\n\n11 AM–2 PM\n\n3\n\nSide course project or practice\n\nSun\n\n9–11 PM\n\n2\n\nPaper summary + 15-min weekly review\n\nThe key idea is running two tracks at once instead of four courses in parallel: one heavy main course (OS, then Networks) and one lighter side course (Go, then Math). Two-hour sessions split across four courses would mean constant context switching and no lab ever getting finished.\n\nSemester 1 week-by-week (26 weeks, with rest weeks at 7, 14, and 21)\n\nWeeks\n\nMain track\n\nSide track\n\n1–6\n\nC1 OS: OSTEP reading + xv6 labs (util + syscall, page tables, traps, COW fork, one lab per week)\n\nWrap up C3 notes and write the Semester 1 blog post\n\nA few reasons behind the ordering: Go comes first because it's short and you'll need it for the MIT 6.5840 labs in Semester 2. The benchmarking toolkit (C3) is finished by week 16 so it's ready for the C2 congestion-control study. OS comes before Networks because the xv6 networking lab and container work make the TCP stack labs easier to reason about. And your networking background means C2 can move faster than its topic list suggests.\n\nPapers for this semester\n\nStart week 1 with Keshav's \"How to Read a Paper,\" then the three datacenter papers the plan assigns to Semester 1 (Al-Fares fat-tree, VL2, Jupiter Rising) in weeks 2–4. After that, pick papers tied to whatever the main track is covering, such as a classic OS or virtualization paper during C1 and a TCP or datacenter congestion paper during C2.\n\nWhen work gets busy\n\nFollow the minimum from your plan: Wednesday paper plus one weekend session (about 4 hours), and never skip two weeks in a row. If a course falls more than a couple of weeks behind this table, drop the lowest-value project part and keep the labs, then use the next rest week as a catch-up week instead of a full break.\n\nWant me to add this as a \"Semester 1 schedule\" tab in your program doc so it sits alongside the course details?\n\nSelf-Directed M.Tech: Networked Systems & Security for AI Infrastructure\n\nSep 30, 2026 · @danisherror\n\nProgram overview\n\nThis is a 2-year, 4-semester self-study program at roughly 15 hours/week (~1,500 hours total), equivalent in depth to an M.Tech in Computer Systems with a research thesis. The goal is to become someone who can build and secure the infrastructure AI runs on: GPU datacenter networks, sandboxed agent execution, and AI-assisted network automation.\n\nWhy this track fits you. It builds on what you already do well (C, control planes, IS-IS/SR, TI-LFA/FRR) instead of restarting in crowded pure ML. It connects directly to your MCP approval system, dev-VM sandbox, Process Network Inspector, and the Apstra work at HPE.\n\nHow it maps to a real M.Tech.\n\nReal M.Tech component\n\nEquivalent here\n\nCore courses (~6)\n\n6 core courses across Semesters 1–3\n\nElectives (~3)\n\n3 electives in Semester 3, chosen from a list\n\nLab / assignments\n\n1–2 graded projects per course\n\nSeminar\n\nPaper-reading practice: 1 paper/week, written summaries\n\nThesis (~1 year)\n\nResearch project from Semester 3, full-time focus in Semester 4\n\nDegree certificate\n\nPublic portfolio: GitHub repos, technical blog, open-source commits, a workshop paper or arXiv preprint\n\nRules that keep it rigorous.\n\nEvery course ends with a project that has a written report (2–4 pages) and a public repo.\n\nUse free university courses (MIT, CMU, Stanford, Berkeley) as the lecture backbone; do their labs, not just the videos.\n\nWrite in C, Go, or Python depending on the layer; C and Go carry most of the systems work.\n\nCourse URLs and editions change; search the course code to find the current offering.\n\nProgram structure\n\n[embedded content: program roadmap · 4 semesters, 3 gates]\n\nSemesters 1 and 2 build the systems base; the thesis starts at month 13 beside the AI-layer courses and takes over fully in Semester 4. Do not pass a gate until its condition is met.\n\nSemester 1 (months 1–6): Systems foundations\n\nSemester 1 closes the gaps under your networking strength: OS internals, modern networking beyond routing, the math research needs, and Go.\n\nC1. Operating Systems & Systems Programming\n\nBackbone: MIT 6.1810 (formerly 6.S081, xv6 labs); book Operating Systems: Three Easy Pieces (free online).\n\nTopics\n\nProcesses, threads, scheduling, context switches\n\nVirtual memory, page tables, TLBs, copy-on-write\n\nSystem calls, traps, interrupts, the kernel/user boundary\n\nMini-container runtime in C/Go: build a tool that uses namespaces, cgroups v2, pivot_root, seccomp, and dropped capabilities to run a process. Write a threat model listing what it does and does not isolate. This formalizes your recent container investigation.\n\nC2. Advanced Computer Networks\n\nBackbone: Stanford CS144 (build a TCP stack); Computer Networks: A Systems Approach by Peterson & Davie (free online); Stanford CS244 or Princeton COS 561 reading lists for advanced papers.\n\nTopics\n\nTCP internals: congestion control (Reno, CUBIC, BBR), flow control, retransmission\n\nNetwork measurement and telemetry: sFlow, INT, gNMI streaming\n\nCongestion in datacenters: ECN, DCTCP, incast\n\nRouting at scale: BGP in the datacenter (RFC 7938), EVPN-VXLAN\n\nProjects\n\nComplete the CS144 TCP implementation labs.\n\nCongestion-control lab: in Mininet or ns-3, compare CUBIC, BBR, and DCTCP under incast and report the results with graphs.\n\nEVPN-VXLAN fabric in containerlab: build a 2-spine/4-leaf fabric with FRRouting, then break links and measure convergence.\n\nC3. Mathematics for Systems Research\n\nBackbone: MIT 6.041 (Probability), Gilbert Strang's MIT 18.06 (Linear Algebra), selected queueing theory chapters (Mor Harchol-Balter, Performance Modeling and Design of Computer Systems).\n\nQueueing theory: M/M/1, Little's law, load vs latency\n\nLinear algebra essentials for ML (matrices, gradients, tensor shapes)\n\nGraph algorithms: shortest paths, max-flow, spanning trees (you know SPF; add flow and cuts)\n\nStatistics for experiments: confidence intervals, variance, how to compare two systems fairly\n\nProject\n\nStatistical benchmarking toolkit (Python): run a workload N times, compute confidence intervals, and draw CDF plots. You will reuse it in every later experiment.\n\nC4. Go for Systems (short course, ~6 weeks)\n\nBackbone:The Go Programming Language (Donovan & Kernighan); Go by Example; Go's official concurrency material.\n\nTopics: goroutines, channels, context cancellation, gRPC/protobuf, profiling with pprof, writing CLIs and daemons.\n\nProject\n\ngNMI telemetry collector in Go: subscribe to streaming telemetry from a containerlab device (SR Linux or cEOS), store it in PostgreSQL or a time-series DB, and alert on link flaps. This maps directly to Apstra device-agent work.\n\nSemester 2 (months 7–12): Core specialization\n\nSemester 2 builds the three pillars of the specialization: distributed systems, networks for AI clusters, and systems security.\n\nC5. Distributed Systems\n\nBackbone: MIT 6.5840 (formerly 6.824) with its Go labs; Designing Data-Intensive Applications (Kleppmann).\n\nReplicated approval store: rebuild the approval/audit store from your MCP dev-VM project on top of your Raft implementation, so approvals survive node failure and the hash-chained audit log stays consistent.\n\nC6. Datacenter Networking for AI Clusters\n\nThis is the course that makes you distinct. No single university course covers it fully, so the backbone is papers plus labs.\n\nBackbone: papers listed under the reading list (Jupiter, fat-tree, DCQCN, RDMA at scale, HPCC, Meta and Alibaba AI-training networks); NVIDIA NCCL documentation; Ultra Ethernet Consortium public material.\n\nTopics\n\nHow distributed training works: data/tensor/pipeline parallelism, all-reduce, all-to-all\n\nCollective communication: ring and tree all-reduce, NCCL internals\n\nFailure handling: link flaps, stragglers, how one failure stalls a training job\n\nFast reroute in the fabric: applying LFA/TI-LFA thinking to AI fabrics\n\nTelemetry and root-cause analysis for training-job slowdowns\n\nProjects\n\nAll-reduce simulator (Python or Go): simulate ring vs tree all-reduce on a fat-tree in ns-3 or a custom flow-level simulator; measure job completion time under one failed link.\n\nECMP collision study: show how a few large flows collide under hash-based ECMP and evaluate a flowlet- or spraying-based fix.\n\nWrite-up: a 3-page survey \"Failure recovery in AI training fabrics\" — this often becomes the seed of the thesis.\n\nC7. Systems Security\n\nBackbone: Stanford CS155 or Berkeley CS161; MIT 6.5660 (formerly 6.858) labs; Ross Anderson, Security Engineering (3rd ed.).\n\nComplete the MIT 6.5660 lab series (privilege separation, web security).\n\nSandbox comparison study: run the same untrusted workload under Docker, gVisor, Firecracker, and Kata on Linux; compare isolation guarantees, startup time, and syscall overhead with your benchmarking toolkit.\n\nFormal threat model for your MCP CRUD system: attack tree, trust boundaries, and a test suite of malicious-agent scenarios.\n\nSemester 3 (months 13–18): AI systems, programmable networks, electives\n\nSemester 3 adds the AI layer on top of your systems base and starts the thesis in parallel (about 5 of the 15 weekly hours go to research from month 13).\n\nC8. ML Systems (how AI actually runs)\n\nBackbone: Andrej Karpathy's Neural Networks: Zero to Hero (build GPT from scratch); CMU 10-414/714 Deep Learning Systems (build a mini PyTorch); Stanford CS336 Language Modeling from Scratch.\n\nTopics\n\nNeural network basics, backpropagation, transformers and attention\n\nTensors, autograd, GPU memory hierarchy, kernels\n\nTraining at scale: data/tensor/pipeline parallelism, ZeRO, checkpointing\n\nBuild a small GPT from scratch in PyTorch and train it on a small dataset (your M2 Mac handles small models; use a cloud GPU for a few hours when needed).\n\nNetwork-trace generator from training: instrument a 2–4 GPU data-parallel job (rented) and capture its all-reduce traffic pattern; feed it into your C6 simulator.\n\nC9. Programmable Networks & Fast Packet Processing\n\nBackbone: p4lang/tutorials on GitHub; Learning eBPF (Liz Rice); DPDK documentation; George Varghese, Network Algorithmics.\n\nComplete the P4 tutorials; implement a P4 program that does per-flow telemetry.\n\neBPF process-network observer on Linux: a Linux counterpart to your Process Network Inspector that maps sockets and flows to PIDs via eBPF, strictly view-only.\n\nFast-reroute in P4: implement data-plane link-failure detection and local reroute in BMv2, and compare failover time against control-plane recomputation.\n\nC10. Electives (pick 3)\n\nElective\n\nBackbone\n\nProject\n\nBest if thesis is on\n\nE1. Security of LLM agents\n\nOWASP Top 10 for LLM Applications; papers on prompt injection and CaMeL; AgentDojo benchmark\n\nRed-team your MCP CRUD system with injected tool outputs; measure how many attacks the approval layer blocks\n\nAgent security\n\nE2. Network verification\n\nBatfish docs; papers on header space analysis and config verification\n\nUse Batfish to verify reachability and ECMP properties of your EVPN fabric before and after a change\n\nSafe AI network automation\n\nE3. Formal methods (TLA+)\n\nLeslie Lamport's TLA+ video course; Specifying Systems\n\nWrite a TLA+ spec of your proposal→approval→revalidation→execute protocol and model-check race conditions\n\nAgent security or safe automation\n\nE4. Intent-based networking\n\nApstra public docs; RFC 9315 (intent-based networking concepts); SONiC docs\n\nBuild a tiny intent engine: YAML intent → rendered FRR config → Batfish check → deploy to containerlab\n\nSafe AI network automation\n\nE5. Performance engineering\n\nBrendan Gregg's books; Systems Performance\n\nProfile and speed up a real open-source network daemon (e.g., FRR isisd) and upstream a patch\n\nAI fabric reliability\n\nE6. Storage systems\n\nCMU 15-445 (databases); papers on distributed file systems\n\nBuild a checkpoint store for training jobs and measure recovery time\n\nAI fabric reliability\n\nRecommended default set: E1 + E2 + E3. They cover security, verification, and formal reasoning, which strengthen any of the thesis topics below.\n\nResearch projects and thesis options\n\nPick one thesis topic by month 12 and work on it from month 13 to month 24. The first option is the strongest match for your background; the other three are solid alternatives.\n\nThesis option A (recommended): Fast failure recovery for AI training fabrics\n\nResearch question: When a link or switch fails in a GPU cluster fabric, how much training time is lost, and can local fast-reroute (in the spirit of TI-LFA) plus collective-aware traffic engineering cut that loss significantly?\n\nWhy you: almost nobody working on AI networking has shipped TI-LFA in production code. You have.\n\nMethod\n\nBuild a flow-level simulator of a rail-optimized Clos fabric driven by real all-reduce traces (from C8).\n\nMeasure baseline: ECMP rehash plus control-plane reconvergence after failures.\n\nDesign a scheme: precomputed backup paths per collective group, placed so the backup avoids congested links.\n\nImplement a prototype in P4 (BMv2) or FRR plus containerlab for the control-plane part.\n\nEvaluate: job completion time, tail latency of collectives, backup-path computation cost, at 64 to 4,096 simulated GPUs.\n\nTarget venues: HotNets, the SIGCOMM or NSDI workshops, APNet; arXiv preprint regardless.\n\nResearch question: Can an LLM agent propose network changes that are provably safe before deployment, using verification (Batfish), formal intent, and human approval gates?\n\nMethod: build a pipeline of natural-language intent → LLM-generated config → Batfish verification → diff-based human approval → staged rollout with automatic rollback. Create a benchmark of 100+ change requests including deliberately risky ones, and measure how many unsafe changes each safety layer catches.\n\nWhy you: combines Apstra-style intent networking with your approval-gated MCP design. Directly useful at HPE.\n\nThesis option C: Verifiable authorization for AI agents acting on real systems\n\nResearch question: How can a system prove that every action an agent took was explicitly authorized by a user, even when the agent is manipulated by prompt injection?\n\nMethod: formalize your proposal → approval → revalidation → execute protocol in TLA+; implement it with canonical proposal hashes and a transparency-log-style audit; attack it with injected tool outputs (AgentDojo-style); measure attack success rate and approval overhead.\n\nWhy you: this is your MCP CRUD and dev-VM work turned into research.\n\nThesis option D: Sandboxing untrusted agent code\n\nResearch question: What is the right isolation boundary (container, gVisor, microVM) for agent-executed code, and what capability model lets agents do useful work without broad access?\n\nMethod: systematically compare escape surface, syscall coverage, and startup latency; design an effect-declaration system (from your dev-VM project) and enforce it with seccomp or eBPF LSM.\n\nSmaller research projects (portfolio pieces)\n\nComplete at least two of these alongside the courses; each is 4–8 weeks.\n\nProject\n\nOutput\n\nLinks to\n\nMeasure ECMP imbalance under AI-style elephant flows\n\nBlog post + repo\n\nC6, thesis A\n\nUpstream a fix or feature to FRRouting isisd (e.g., TI-LFA improvements)\n\nMerged patch\n\nYour IS-IS-SR expertise\n\nLLM agent that diagnoses a broken containerlab fabric from telemetry\n\nDemo + write-up\n\nC4, C8, thesis B\n\nPrompt-injection test suite for MCP servers\n\nOpen-source tool\n\nE1, thesis C\n\nTLA+ spec of an approval protocol, with a found bug documented\n\nBlog post\n\nE3, thesis C\n\nPaper reading list and research skills\n\nRead one paper a week (~100 over two years) and write a 1-page summary for each: problem, key idea, evaluation, weakness, and one follow-up idea. Start with S. Keshav's short guide How to Read a Paper (three-pass method).\n\nWhere the field publishes: SIGCOMM, NSDI, HotNets, CoNEXT (networking); OSDI, SOSP, EuroSys, USENIX ATC (systems); USENIX Security, IEEE S&P, CCS (security); MLSys (ML systems).\n\nArea\n\nPapers to start with\n\nRead in\n\nDatacenter networks\n\nA Scalable, Commodity Data Center Network Architecture (Al-Fares et al., 2008); VL2 (2009); Jupiter Rising (Google, 2015)\n\nSem 1\n\nLossless / RDMA networking\n\nCongestion Control for Large-Scale RDMA Deployments (DCQCN, 2015); RDMA over Commodity Ethernet at Scale (2016); HPCC (2019); Swift (2020)\n\nSem 2\n\nAI training networks\n\nRDMA over Ethernet for Distributed AI Training at Meta Scale (2024); Alibaba HPN (2024)\n\nSem 2\n\nDistributed systems\n\nRaft (2014); Paxos Made Simple; MapReduce; GFS; Spanner; Borg\n\nHeader Space Analysis (2012); A General Approach to Network Configuration Verification (Batfish, 2017)\n\nSem 3\n\nIsolation\n\nFirecracker (NSDI 2020); gVisor design docs\n\nSem 2\n\nML systems\n\nMegatron-LM; ZeRO; Efficient Memory Management for LLM Serving with PagedAttention (vLLM, 2023)\n\nSem 3\n\nAgent security\n\nNot what you've signed up for: indirect prompt injection (Greshake et al., 2023); AgentDojo (2024); Defeating Prompt Injections by Design (CaMeL, 2025)\n\nSem 3\n\nResearch skills to practice (spread over all four semesters)\n\nFinding a gap: for each paper, write what it did not evaluate.\n\nWriting: follow the structure of a 6-page workshop paper (problem, motivation with numbers, design, evaluation, related work).\n\nFigures: CDFs and time series with clearly labeled axes; reuse your benchmarking toolkit.\n\nGetting feedback: email authors with specific questions, post drafts publicly, ask for reviews in research communities.\n\nWeekly schedule (fits your three-block day)\n\nThe program needs about 15 hours a week, taken from your 6 PM – 3 AM personal block: 9 hours on weekdays and 6 on weekends. Times are flexible; the weekly hour count is what matters.\n\nDay\n\nTime (IST)\n\nHours\n\nActivity\n\nMonday\n\n9 – 11 PM\n\n2\n\nLectures and reading for the current course\n\nTuesday\n\n9 – 11 PM\n\n2\n\nLab / project coding\n\nWednesday\n\n9 – 10 PM\n\n1\n\nPaper of the week (first and second pass)\n\nThursday\n\n9 – 11 PM\n\n2\n\nLab / project coding\n\nFriday\n\n—\n\n0\n\nRest; buffer for work overruns\n\nSaturday\n\n11 AM – 2 PM\n\n3\n\nDeep work: hardest lab or thesis experiments\n\nSunday\n\n11 AM – 2 PM, 9 – 11 PM\n\n3 + 2\n\nDeep work; then paper summary + weekly review\n\nAdjustments\n\nIn release crunches at work, drop to the minimum: Wednesday paper + one weekend session (~4 hours). Never skip two weeks in a row.\n\nFrom month 13, Saturday becomes a thesis day.\n\nEvery 6 weeks, take one full week off. That keeps two years sustainable.\n\nWeekly review (15 minutes, Sunday night)\n\nWhat did I finish this week?\n\nWhat is blocked, and why?\n\nIs the current course on track for its 6-month window?\n\nNext week's single most important task\n\nAssessment, milestones, and credibility without a degree\n\nWithout a university, your proof is public work. Aim for evidence that a hiring manager or PhD advisor can check in five minutes.\n\nHow to grade yourself (per course)\n\nComponent\n\nWeight\n\nPass bar\n\nLabs completed\n\n40%\n\nAll official labs pass their test suites\n\nCourse project\n\n40%\n\nWorking repo + 2–4 page report with measured results\n\nPaper summaries\n\n10%\n\nOne per week, no gaps longer than 2 weeks\n\nPublic write-up\n\n10%\n\nOne blog post explaining what you learned\n\nEnd-of-program portfolio checklist\n\n10 course project repos with READMEs and reports\n\n1 thesis (40–60 pages) and its code, public on GitHub\n\n1 arXiv preprint or workshop paper submission from the thesis\n\nAt least 1 merged patch in a real project (FRRouting, containerlab, Batfish, gVisor, or similar)\n\nSelf-Directed Research Program: MS by Research and PhD Track\n\nSep 30, 2026 · @danisherror\n\nProgram overview\n\nThis program trains you to produce publishable research, not to finish courses. It has two tracks: an MS by Research equivalent (about 2.5 years, 1–2 papers and a thesis) and a PhD equivalent (about 5 years total, 3–5 papers and a dissertation). The PhD track continues from the MS track; you decide at the MS gate whether to go on.\n\nHow it differs from the M.Tech plan\n\nM.Tech plan\n\nThis research program\n\nMain output\n\nSkills + projects\n\nNew knowledge: peer-reviewed papers\n\nCoursework\n\n~70% of time\n\n~20%, front-loaded in year 1\n\nSuccess measure\n\nLabs pass, projects work\n\nReviewers at SIGCOMM/NSDI/USENIX accept the work\n\nTime budget\n\n~15 h/week\n\n~18–20 h/week (research needs long uninterrupted blocks)\n\nResearch area:Dependable infrastructure for AI — networks and agent-execution systems that fail safely. It combines your control-plane depth (TI-LFA, IS-IS-SR) with your agent-security work (approval-gated MCP, sandboxing).\n\nAn honest limit. Self-study cannot grant a degree, and research without feedback drifts. This program therefore builds in external reviewers, collaborators, and real paper submissions from year 1, and ends with a path to register the work for a formal degree (last section).\n\nBoth tracks share the first 30 months; at Gate 2 you either stop with a completed MS thesis or continue into the PhD phases.\n\nPhase 1 (months 1–10): Research foundations\n\nPhase 1 replaces the coursework year of a research degree: four research-skill modules, a breadth survey, and a qualifying exam you set for yourself and have someone else grade.\n\nR1. Reading and critiquing research (months 1–3)\n\nMethod: S. Keshav's three-pass reading; for each paper, write a review in conference format (summary, strengths, weaknesses, questions to authors, accept/reject).\n\nVolume: 3 papers/week for 12 weeks (~36 reviews).\n\nCalibration: pick papers from venues that publish reviews or public discussion (e.g., OpenReview) and compare your review with the real ones.\n\nOutput: a review notebook on GitHub.\n\nR2. Experimental methodology (months 2–5)\n\nTopics\n\nChoosing baselines that a reviewer will accept\n\nWorkloads: real traces vs synthetic; where to find public datacenter and training-job traces\n\nStatistics: confidence intervals, repeated runs, variance sources on shared machines\n\nSimulation vs emulation vs testbed: ns-3, containerlab/Mininet, real hardware; validating a simulator against reality\n\nReproducibility: artifact packaging, scripts that regenerate every figure\n\nProject: reproduce one published result (for example, a DCQCN or ECMP-imbalance figure) in simulation and write up where your numbers differ and why.\n\nR3. Research writing (months 4–10, continuous)\n\nStudy: Simon Peyton Jones's talk How to Write a Great Research Paper; Jim Kurose's advice on writing paper introductions.\n\nPractice: rewrite the introduction of your reproduction report five times, getting feedback each time.\n\nLearn LaTeX and the ACM/USENIX templates; keep a figure-generation pipeline in Python.\n\nR4. Breadth survey (months 5–9)\n\nWrite a 15–20 page survey: \"Failure handling in AI infrastructure: networks, collectives, and agent execution\". Cover ~60–80 papers, organize them into a taxonomy, and end with 5–10 open problems. This becomes your map for the next four years and can be posted on arXiv.\n\nQualifying exam equivalent (month 10)\n\nPart\n\nFormat\n\nWho grades it\n\nWritten\n\nAnswer 3 of 5 questions on your survey area, 48 hours, open book\n\nAn external mentor (see the advisors section)\n\nOral\n\n45-minute talk on your survey + open problems, then 30 minutes of questions\n\nMentor + 1–2 engineers or researchers\n\nPass bar\n\nYou can defend why each open problem matters and how you would evaluate a solution\n\nMentor's judgment\n\nIf you fail, spend 6 more weeks on the weak area and retake it. Do not start the main research agenda before passing.\n\nResearch agenda\n\nCentral question: How do we build AI infrastructure whose failures, whether hardware faults in the network or wrong actions by an AI agent, are contained quickly and provably, without a human debugging every incident?\n\nA PhD is a chain of papers that each answer one part of this question. The first two papers form the MS thesis; papers 3–5 extend it into a PhD.\n\n#\n\nPaper (working title)\n\nResearch question\n\nMethod\n\nTarget venue\n\nTrack\n\nP1\n\nMeasuring the cost of link failures in AI training fabrics\n\nHow much job time does one link or switch failure cost at different scales, and where does it go (detection, reconvergence, stragglers)?\n\nTrace-driven simulation of rail-optimized Clos fabrics; validation on a small rented GPU cluster\n\nHotNets or APNet (workshop)\n\nMS\n\nP2\n\nCollective-aware fast reroute for GPU fabrics\n\nCan precomputed, collective-aware backup paths (TI-LFA-style) cut the time lost to failures?\n\nCan every agent action be tied to an explicit, revalidated user approval, even under prompt injection?\n\nTLA+ protocol spec, implementation from your MCP work, adversarial evaluation\n\nUSENIX Security or CCS\n\nPhD\n\nP5\n\nUnified containment for AI infrastructure\n\nCan one framework cover both network faults and agent faults (detect, contain, recover, audit)?\n\nSynthesis of P2–P4 into one system; end-to-end evaluation\n\nOSDI or NSDI\n\nPhD\n\nWhy this chain works: P1 produces the numbers that motivate P2. P3 and P4 reuse the approval and verification ideas from each other. P5 is the dissertation's unifying contribution. Each paper can stand alone if a later one fails.\n\nKeep a backup topic. If P1 shows that failure cost is already small in practice, pivot to P3/P4 (agent safety) as the main line; the survey from Phase 1 will show other open problems.\n\nMS by Research track (months 11–30)\n\nThe MS track ends with a thesis built on P1 and P2, at least one submitted paper, and a public defense.\n\nStages\n\nMonths 11–12: Research proposal (5–8 pages). Problem, why it matters (with numbers from your survey), related work, planned approach, evaluation plan, risks. Your mentor must approve it.\n\nMonths 13–18: P1 measurement study. Build the simulator, validate it on a small real setup, publish the tool as open source. Submit P1 to a workshop.\n\nMonths 19–27: P2 system. Design the fast-reroute scheme, prototype it, evaluate at scale. Submit P2 to a main conference (expect rejection on the first try; revise and resubmit).\n\nMonths 28–30: Thesis writing and defense. Write the thesis and give a 45-minute public defense (online meetup, recorded) with your mentor and two outside reviewers asking questions.\n\nBackground: AI training traffic, datacenter fabrics, fast reroute (LFA, RLFA, TI-LFA)\n\nMeasurement study (P1)\n\nCollective-aware fast reroute: design (P2)\n\nImplementation\n\nEvaluation\n\nRelated work\n\nLimitations and future work\n\nAppendix: artifact and reproduction instructions\n\nMS exit criteria\n\nProposal approved by mentor\n\nP1 submitted (workshop) and posted on arXiv\n\nP2 submitted to a main venue at least once\n\nCode and data public, with a script that regenerates every figure\n\nThesis written and defended\n\nDecision at month 30: continue to the PhD track only if you enjoyed the uncertainty of research, P2 got substantive reviews (even if rejected), and you still have a question you want to answer. Otherwise, the MS is a strong finish on its own.\n\nPhD track (months 31–60)\n\nThe PhD track adds P3, P4, and P5 on top of the MS work and ends with a dissertation whose central claim ties them together.\n\nStages\n\nMonths 31–33: Dissertation proposal (15–20 pages). Thesis statement, completed work (P1, P2), planned work (P3–P5), timeline, and risks. Defend it orally before a committee of three: your mentor plus two outside researchers.\n\nMonths 34–42: P3, verified AI-generated network change. Build the pipeline and a public benchmark of change requests; submit to NSDI or CoNEXT.\n\nMonths 40–50: P4, provable agent authorization. Formal spec in TLA+, implementation, adversarial evaluation; submit to USENIX Security or CCS. Overlaps P3 by design so that a rejection cycle on one keeps the other moving.\n\nMonths 50–56: P5, unified containment. The capstone system.\n\nMonths 56–60: Dissertation and defense.\n\nDissertation structure (120–180 pages)\n\nIntroduction and thesis statement, e.g.: \"Failures in AI infrastructure, whether in the network or in agent actions, can be contained within bounded time and with verifiable authorization by combining precomputed local recovery, pre-deployment verification, and explicit approval protocols.\"\n\nBackground and survey (from R4, updated)\n\nPart I: Network failures (P1, P2)\n\nPart II: Agent and automation failures (P3, P4)\n\nPart III: Unified containment (P5)\n\nRelated work, limitations, future directions\n\nPublication targets\n\nLevel\n\nCount by month 60\n\nMinimum\n\n3 papers accepted, at least 1 at a top venue (SIGCOMM, NSDI, OSDI, SOSP, USENIX Security, IEEE S&P, CCS)\n\nStrong\n\n4–5 accepted, 2+ at top venues, one artifact badge (e.g., artifact evaluation at USENIX)\n\nAlso expected\n\nServe as a reviewer or shadow PC member at least once; 2+ public talks\n\nRejections are normal: many top-venue papers are rejected once or twice before acceptance. Budget for one resubmission per paper in the timeline.\n\nAdvisors, collaborators, compute, and testbeds\n\nThe single biggest risk of independent research is working alone. Secure an external mentor by month 6 and a co-author by month 18.\n\nFinding a mentor (the advisor substitute)\n\nFrom your survey, list 15–20 researchers (faculty at IISc, IITs, or abroad; researchers at Microsoft Research India, Google, Meta, NVIDIA) whose papers you cite most.\n\nEmail them with something concrete: a reproduction result that differs from their paper, a bug in their artifact, or a short question about their evaluation. Never send a generic \"please mentor me\" email.\n\nAsk for a light commitment: a 30-minute call once a month.\n\nExpect a low response rate; one yes out of twenty is a good outcome.\n\nOther sources of feedback\n\nIndustry colleagues at HPE with networking depth: ask them to review drafts and serve on your \"committee\".\n\nResearch communities and workshops: submit to workshops and student/poster sessions (many systems conferences have them) for early feedback.\n\nOpen-source maintainers: FRRouting, containerlab, and Batfish developers can review the practical side of your work.\n\nApply to research programs that take working professionals as collaborators or visiting researchers; check what Microsoft Research India and IISc/IIT labs currently offer.\n\nCompute and testbeds\n\nNeed\n\nOption\n\nNotes\n\nNetwork emulation\n\ncontainerlab, Mininet, BMv2 on your Mac or a Linux VM\n\nFree; enough for P2, P3 prototypes\n\nLarge-scale simulation\n\nns-3 or a custom flow-level simulator on a rented multi-core cloud VM\n\nPay per hour; keep runs scripted\n\nReal GPU traffic traces\n\nA few hours on rented 4–8 GPU cloud instances\n\nBudget carefully; capture traces once and reuse them\n\nResearch testbeds\n\nCloudLab, FABRIC, Chameleon (US research testbeds)\n\nAccess usually requires an academic affiliation; a collaborator can give you access\n\nCloud credits\n\nResearch credit programs from cloud providers\n\nUsually need a proposal; your research proposal doubles as the application\n\nBudget: plan roughly for cloud spending of a few thousand rupees a month, rising during evaluation phases. Collaborating with an academic lab is the best way to cut this.\n\nWeekly routine and year-end checkpoints\n\nResearch needs long, uninterrupted blocks, so this routine puts most hours on weekends and keeps weekdays for reading, writing, and small tasks. Total: about 18–20 hours/week.\n\nDay\n\nTime (IST)\n\nHours\n\nActivity\n\nMonday\n\n9 – 11 PM\n\n2\n\nRead: 1 paper, write its review\n\nTuesday\n\n9 – 11 PM\n\n2\n\nCode: small experiment changes, bug fixes\n\nWednesday\n\n9 – 10:30 PM\n\n1.5\n\nWrite: 500 words on the current paper or thesis\n\nThursday\n\n9 – 11 PM\n\n2\n\nCode or analyze last weekend's results\n\nFriday\n\n—\n\n0\n\nRest\n\nSaturday\n\n10 AM – 3 PM\n\n5\n\nDeep work: design, large experiments\n\nSunday\n\n10 AM – 2 PM, 9 – 11 PM\n\n4 + 2\n\nDeep work; then research log + plan next week\n\nHabits that matter more than hours\n\nKeep a dated research log: every experiment, its hypothesis, the result, and what you concluded. This becomes your thesis's raw material.\n\nStart long simulations on Sunday night so they finish during the week.\n\nMonthly: a 30-minute mentor call with a 1-page update sent two days before.\n\nBefore each paper deadline, take 1–2 weeks of leave from work if possible; the last weeks before submission are the most intense.\n\nYear-end checkpoints\n\nEnd of year 1: qualifier passed, survey on arXiv, mentor secured, research proposal approved\n\nEnd of year 2: P1 submitted, P2 prototype working with first results\n\nEnd of year 2.5: MS thesis defended; continue-or-stop decision made\n\nEnd of year 3: dissertation proposal defended, P3 submitted\n\nEnd of year 4: P4 submitted, at least 2 papers accepted overall\n\nEnd of year 5: P5 done, dissertation defended\n\nEvaluating progress and turning it into a formal degree\n\nPeer review is the real grade: a paper accepted at a top venue is the same evidence whether or not you are enrolled anywhere. Track these signals every 6 months.\n\nSignal\n\nHealthy\n\nWarning\n\nResearch log\n\nEntries most weeks\n\nGaps of a month or more\n\nSubmissions\n\nOne submission every 6–9 months after year 1\n\nNothing submitted in 12 months\n\nReviews received\n\nReviewers engage with the idea, criticize the evaluation\n\nReviewers say the problem is not important\n\nMentor contact\n\nMonthly calls happening\n\nNo contact in 3 months\n\nYour motivation\n\nYou think about the problem outside scheduled hours\n\nOnly deadlines make you work\n\nIf reviewers repeatedly say the problem is not important, revisit the agenda with your mentor instead of polishing the same paper.\n\nPaths to a formal degree later\n\nYour published work makes these routes much stronger. Rules and eligibility change, so check each institute's current admission pages.\n\nPart-time or external PhD at IITs/IISc for working professionals: some programs let employees register while working, usually with a short coursework requirement and sometimes an employer sponsorship letter. Your papers and a willing faculty co-author make admission far more likely.\n\nPhD by publication: some universities (mainly in the UK and parts of Europe and Australia) award a PhD based on a portfolio of published papers plus a written commentary; check eligibility for external candidates.\n\nIndustry research roles: papers at NSDI/SIGCOMM/USENIX Security can qualify you for research-engineer roles at industry labs even without a PhD.\n\nFull-time admission later: if your situation changes, a strong publication record also helps for a funded full-time PhD in India or abroad.\n\nThe mentor you find in year 1 matters here too: a faculty member who has co-authored with you is the most natural PhD advisor if you register formally.", "url": "https://wpnews.pro/news/0-mtech-phd-md", "canonical_source": "https://gist.github.com/danisherror/244c1eafc1b0814f50d6357af5134b94", "published_at": "2026-09-30 12:03:44+00:00", "updated_at": "2026-09-30 12:19:34.844960+00:00", "lang": "en", "topics": ["ai-infrastructure", "ai-agents", "mlops"], "entities": ["@danisherror", "HPE", "MIT 6.5840", "xv6", "OSTEP", "Apstra", "MCP", "Keshav"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/0-mtech-phd-md", "markdown": "https://wpnews.pro/news/0-mtech-phd-md.md", "text": "https://wpnews.pro/news/0-mtech-phd-md.txt", "jsonld": "https://wpnews.pro/news/0-mtech-phd-md.jsonld"}}