[0] Mtech_phd.md A developer known as @danisherror has published a self-directed M.Tech program in Networked Systems & Security for AI Infrastructure, a two-year, four-semester self-study curriculum of roughly 15 hours per week (~1,500 hours total) aimed at building and securing the infrastructure AI runs on, including GPU datacenter networks, sandboxed agent execution, and AI-assisted network automation. The plan sequences six core courses across three semesters plus three electives, running a heavy main track (operating systems, then networks) alongside a lighter side track (Go, then math), with xv6 labs, MIT 6.5840 material, and datacenter papers such as Al-Fares fat-tree, VL2, and Jupiter Rising. It also specifies a weekly schedule, rest weeks every six weeks, and a fallback minimum of about four hours per week when work gets busy. Got it, this is your self-directed M.Tech in Networked Systems & Security for AI Infrastructure. That changes things from my earlier generic answer, so here's a schedule built specifically around this program's Semester 1 C1 OS, C2 Advanced Networks, C3 Math, C4 Go , fitted to a 5-hour workday. How much time to use Your plan is designed for about 15 hours a week. With a 5-hour job you have room for more, but I'd keep 15 hours as the committed load and treat the extra capacity as a buffer Friday catch-up, longer lab sessions when something breaks . Two years is long, and the plan's own rule of a week off every 6 weeks only works if the weekly load stays sustainable. The weekly routine same slots your doc uses; shift them if your work hours differ Day Time Hrs What Mon 9–11 PM 2 Main course: lectures and reading Tue 9–11 PM 2 Main course: lab coding Wed 9–10 PM 1 Paper of the week first and second pass Thu 9–11 PM 2 Side course: lectures, exercises Fri 9–11 PM 0–2 Buffer only: catch up if behind, otherwise rest Sat 11 AM–2 PM 3 Deep work: main course lab or project Sun 11 AM–2 PM 3 Side course project or practice Sun 9–11 PM 2 Paper summary + 15-min weekly review The key idea is running two tracks at once instead of four courses in parallel: one heavy main course OS, then Networks and one lighter side course Go, then Math . Two-hour sessions split across four courses would mean constant context switching and no lab ever getting finished. Semester 1 week-by-week 26 weeks, with rest weeks at 7, 14, and 21 Weeks Main track Side track 1–6 C1 OS: OSTEP reading + xv6 labs util + syscall, page tables, traps, COW fork, one lab per week Wrap up C3 notes and write the Semester 1 blog post A few reasons behind the ordering: Go comes first because it's short and you'll need it for the MIT 6.5840 labs in Semester 2. The benchmarking toolkit C3 is finished by week 16 so it's ready for the C2 congestion-control study. OS comes before Networks because the xv6 networking lab and container work make the TCP stack labs easier to reason about. And your networking background means C2 can move faster than its topic list suggests. Papers for this semester Start week 1 with Keshav's "How to Read a Paper," then the three datacenter papers the plan assigns to Semester 1 Al-Fares fat-tree, VL2, Jupiter Rising in weeks 2–4. After that, pick papers tied to whatever the main track is covering, such as a classic OS or virtualization paper during C1 and a TCP or datacenter congestion paper during C2. When work gets busy Follow the minimum from your plan: Wednesday paper plus one weekend session about 4 hours , and never skip two weeks in a row. If a course falls more than a couple of weeks behind this table, drop the lowest-value project part and keep the labs, then use the next rest week as a catch-up week instead of a full break. Want me to add this as a "Semester 1 schedule" tab in your program doc so it sits alongside the course details? Self-Directed M.Tech: Networked Systems & Security for AI Infrastructure Sep 30, 2026 · @danisherror Program overview This is a 2-year, 4-semester self-study program at roughly 15 hours/week ~1,500 hours total , equivalent in depth to an M.Tech in Computer Systems with a research thesis. The goal is to become someone who can build and secure the infrastructure AI runs on: GPU datacenter networks, sandboxed agent execution, and AI-assisted network automation. Why this track fits you. It builds on what you already do well C, control planes, IS-IS/SR, TI-LFA/FRR instead of restarting in crowded pure ML. It connects directly to your MCP approval system, dev-VM sandbox, Process Network Inspector, and the Apstra work at HPE. How it maps to a real M.Tech. Real M.Tech component Equivalent here Core courses ~6 6 core courses across Semesters 1–3 Electives ~3 3 electives in Semester 3, chosen from a list Lab / assignments 1–2 graded projects per course Seminar Paper-reading practice: 1 paper/week, written summaries Thesis ~1 year Research project from Semester 3, full-time focus in Semester 4 Degree certificate Public portfolio: GitHub repos, technical blog, open-source commits, a workshop paper or arXiv preprint Rules that keep it rigorous. Every course ends with a project that has a written report 2–4 pages and a public repo. Use free university courses MIT, CMU, Stanford, Berkeley as the lecture backbone; do their labs, not just the videos. Write in C, Go, or Python depending on the layer; C and Go carry most of the systems work. Course URLs and editions change; search the course code to find the current offering. Program structure embedded content: program roadmap · 4 semesters, 3 gates Semesters 1 and 2 build the systems base; the thesis starts at month 13 beside the AI-layer courses and takes over fully in Semester 4. Do not pass a gate until its condition is met. Semester 1 months 1–6 : Systems foundations Semester 1 closes the gaps under your networking strength: OS internals, modern networking beyond routing, the math research needs, and Go. C1. Operating Systems & Systems Programming Backbone: MIT 6.1810 formerly 6.S081, xv6 labs ; book Operating Systems: Three Easy Pieces free online . Topics Processes, threads, scheduling, context switches Virtual memory, page tables, TLBs, copy-on-write System calls, traps, interrupts, the kernel/user boundary Mini-container runtime in C/Go: build a tool that uses namespaces, cgroups v2, pivot root, seccomp, and dropped capabilities to run a process. Write a threat model listing what it does and does not isolate. This formalizes your recent container investigation. C2. Advanced Computer Networks Backbone: Stanford CS144 build a TCP stack ; Computer Networks: A Systems Approach by Peterson & Davie free online ; Stanford CS244 or Princeton COS 561 reading lists for advanced papers. Topics TCP internals: congestion control Reno, CUBIC, BBR , flow control, retransmission Network measurement and telemetry: sFlow, INT, gNMI streaming Congestion in datacenters: ECN, DCTCP, incast Routing at scale: BGP in the datacenter RFC 7938 , EVPN-VXLAN Projects Complete the CS144 TCP implementation labs. Congestion-control lab: in Mininet or ns-3, compare CUBIC, BBR, and DCTCP under incast and report the results with graphs. EVPN-VXLAN fabric in containerlab: build a 2-spine/4-leaf fabric with FRRouting, then break links and measure convergence. C3. Mathematics for Systems Research Backbone: MIT 6.041 Probability , Gilbert Strang's MIT 18.06 Linear Algebra , selected queueing theory chapters Mor Harchol-Balter, Performance Modeling and Design of Computer Systems . Queueing theory: M/M/1, Little's law, load vs latency Linear algebra essentials for ML matrices, gradients, tensor shapes Graph algorithms: shortest paths, max-flow, spanning trees you know SPF; add flow and cuts Statistics for experiments: confidence intervals, variance, how to compare two systems fairly Project Statistical benchmarking toolkit Python : run a workload N times, compute confidence intervals, and draw CDF plots. You will reuse it in every later experiment. C4. Go for Systems short course, ~6 weeks Backbone:The Go Programming Language Donovan & Kernighan ; Go by Example; Go's official concurrency material. Topics: goroutines, channels, context cancellation, gRPC/protobuf, profiling with pprof, writing CLIs and daemons. Project gNMI telemetry collector in Go: subscribe to streaming telemetry from a containerlab device SR Linux or cEOS , store it in PostgreSQL or a time-series DB, and alert on link flaps. This maps directly to Apstra device-agent work. Semester 2 months 7–12 : Core specialization Semester 2 builds the three pillars of the specialization: distributed systems, networks for AI clusters, and systems security. C5. Distributed Systems Backbone: MIT 6.5840 formerly 6.824 with its Go labs; Designing Data-Intensive Applications Kleppmann . Replicated approval store: rebuild the approval/audit store from your MCP dev-VM project on top of your Raft implementation, so approvals survive node failure and the hash-chained audit log stays consistent. C6. Datacenter Networking for AI Clusters This is the course that makes you distinct. No single university course covers it fully, so the backbone is papers plus labs. Backbone: papers listed under the reading list Jupiter, fat-tree, DCQCN, RDMA at scale, HPCC, Meta and Alibaba AI-training networks ; NVIDIA NCCL documentation; Ultra Ethernet Consortium public material. Topics How distributed training works: data/tensor/pipeline parallelism, all-reduce, all-to-all Collective communication: ring and tree all-reduce, NCCL internals Failure handling: link flaps, stragglers, how one failure stalls a training job Fast reroute in the fabric: applying LFA/TI-LFA thinking to AI fabrics Telemetry and root-cause analysis for training-job slowdowns Projects All-reduce simulator Python or Go : simulate ring vs tree all-reduce on a fat-tree in ns-3 or a custom flow-level simulator; measure job completion time under one failed link. ECMP collision study: show how a few large flows collide under hash-based ECMP and evaluate a flowlet- or spraying-based fix. Write-up: a 3-page survey "Failure recovery in AI training fabrics" — this often becomes the seed of the thesis. C7. Systems Security Backbone: Stanford CS155 or Berkeley CS161; MIT 6.5660 formerly 6.858 labs; Ross Anderson, Security Engineering 3rd ed. . Complete the MIT 6.5660 lab series privilege separation, web security . Sandbox comparison study: run the same untrusted workload under Docker, gVisor, Firecracker, and Kata on Linux; compare isolation guarantees, startup time, and syscall overhead with your benchmarking toolkit. Formal threat model for your MCP CRUD system: attack tree, trust boundaries, and a test suite of malicious-agent scenarios. Semester 3 months 13–18 : AI systems, programmable networks, electives Semester 3 adds the AI layer on top of your systems base and starts the thesis in parallel about 5 of the 15 weekly hours go to research from month 13 . C8. ML Systems how AI actually runs Backbone: Andrej Karpathy's Neural Networks: Zero to Hero build GPT from scratch ; CMU 10-414/714 Deep Learning Systems build a mini PyTorch ; Stanford CS336 Language Modeling from Scratch. Topics Neural network basics, backpropagation, transformers and attention Tensors, autograd, GPU memory hierarchy, kernels Training at scale: data/tensor/pipeline parallelism, ZeRO, checkpointing Build a small GPT from scratch in PyTorch and train it on a small dataset your M2 Mac handles small models; use a cloud GPU for a few hours when needed . Network-trace generator from training: instrument a 2–4 GPU data-parallel job rented and capture its all-reduce traffic pattern; feed it into your C6 simulator. C9. Programmable Networks & Fast Packet Processing Backbone: p4lang/tutorials on GitHub; Learning eBPF Liz Rice ; DPDK documentation; George Varghese, Network Algorithmics. Complete the P4 tutorials; implement a P4 program that does per-flow telemetry. eBPF process-network observer on Linux: a Linux counterpart to your Process Network Inspector that maps sockets and flows to PIDs via eBPF, strictly view-only. Fast-reroute in P4: implement data-plane link-failure detection and local reroute in BMv2, and compare failover time against control-plane recomputation. C10. Electives pick 3 Elective Backbone Project Best if thesis is on E1. Security of LLM agents OWASP Top 10 for LLM Applications; papers on prompt injection and CaMeL; AgentDojo benchmark Red-team your MCP CRUD system with injected tool outputs; measure how many attacks the approval layer blocks Agent security E2. Network verification Batfish docs; papers on header space analysis and config verification Use Batfish to verify reachability and ECMP properties of your EVPN fabric before and after a change Safe AI network automation E3. Formal methods TLA+ Leslie Lamport's TLA+ video course; Specifying Systems Write a TLA+ spec of your proposal→approval→revalidation→execute protocol and model-check race conditions Agent security or safe automation E4. Intent-based networking Apstra public docs; RFC 9315 intent-based networking concepts ; SONiC docs Build a tiny intent engine: YAML intent → rendered FRR config → Batfish check → deploy to containerlab Safe AI network automation E5. Performance engineering Brendan Gregg's books; Systems Performance Profile and speed up a real open-source network daemon e.g., FRR isisd and upstream a patch AI fabric reliability E6. Storage systems CMU 15-445 databases ; papers on distributed file systems Build a checkpoint store for training jobs and measure recovery time AI fabric reliability Recommended default set: E1 + E2 + E3. They cover security, verification, and formal reasoning, which strengthen any of the thesis topics below. Research projects and thesis options Pick one thesis topic by month 12 and work on it from month 13 to month 24. The first option is the strongest match for your background; the other three are solid alternatives. Thesis option A recommended : Fast failure recovery for AI training fabrics Research question: When a link or switch fails in a GPU cluster fabric, how much training time is lost, and can local fast-reroute in the spirit of TI-LFA plus collective-aware traffic engineering cut that loss significantly? Why you: almost nobody working on AI networking has shipped TI-LFA in production code. You have. Method Build a flow-level simulator of a rail-optimized Clos fabric driven by real all-reduce traces from C8 . Measure baseline: ECMP rehash plus control-plane reconvergence after failures. Design a scheme: precomputed backup paths per collective group, placed so the backup avoids congested links. Implement a prototype in P4 BMv2 or FRR plus containerlab for the control-plane part. Evaluate: job completion time, tail latency of collectives, backup-path computation cost, at 64 to 4,096 simulated GPUs. Target venues: HotNets, the SIGCOMM or NSDI workshops, APNet; arXiv preprint regardless. Research question: Can an LLM agent propose network changes that are provably safe before deployment, using verification Batfish , formal intent, and human approval gates? Method: build a pipeline of natural-language intent → LLM-generated config → Batfish verification → diff-based human approval → staged rollout with automatic rollback. Create a benchmark of 100+ change requests including deliberately risky ones, and measure how many unsafe changes each safety layer catches. Why you: combines Apstra-style intent networking with your approval-gated MCP design. Directly useful at HPE. Thesis option C: Verifiable authorization for AI agents acting on real systems Research question: How can a system prove that every action an agent took was explicitly authorized by a user, even when the agent is manipulated by prompt injection? Method: formalize your proposal → approval → revalidation → execute protocol in TLA+; implement it with canonical proposal hashes and a transparency-log-style audit; attack it with injected tool outputs AgentDojo-style ; measure attack success rate and approval overhead. Why you: this is your MCP CRUD and dev-VM work turned into research. Thesis option D: Sandboxing untrusted agent code Research question: What is the right isolation boundary container, gVisor, microVM for agent-executed code, and what capability model lets agents do useful work without broad access? Method: systematically compare escape surface, syscall coverage, and startup latency; design an effect-declaration system from your dev-VM project and enforce it with seccomp or eBPF LSM. Smaller research projects portfolio pieces Complete at least two of these alongside the courses; each is 4–8 weeks. Project Output Links to Measure ECMP imbalance under AI-style elephant flows Blog post + repo C6, thesis A Upstream a fix or feature to FRRouting isisd e.g., TI-LFA improvements Merged patch Your IS-IS-SR expertise LLM agent that diagnoses a broken containerlab fabric from telemetry Demo + write-up C4, C8, thesis B Prompt-injection test suite for MCP servers Open-source tool E1, thesis C TLA+ spec of an approval protocol, with a found bug documented Blog post E3, thesis C Paper reading list and research skills Read one paper a week ~100 over two years and write a 1-page summary for each: problem, key idea, evaluation, weakness, and one follow-up idea. Start with S. Keshav's short guide How to Read a Paper three-pass method . Where the field publishes: SIGCOMM, NSDI, HotNets, CoNEXT networking ; OSDI, SOSP, EuroSys, USENIX ATC systems ; USENIX Security, IEEE S&P, CCS security ; MLSys ML systems . Area Papers to start with Read in Datacenter networks A Scalable, Commodity Data Center Network Architecture Al-Fares et al., 2008 ; VL2 2009 ; Jupiter Rising Google, 2015 Sem 1 Lossless / RDMA networking Congestion Control for Large-Scale RDMA Deployments DCQCN, 2015 ; RDMA over Commodity Ethernet at Scale 2016 ; HPCC 2019 ; Swift 2020 Sem 2 AI training networks RDMA over Ethernet for Distributed AI Training at Meta Scale 2024 ; Alibaba HPN 2024 Sem 2 Distributed systems Raft 2014 ; Paxos Made Simple; MapReduce; GFS; Spanner; Borg Header Space Analysis 2012 ; A General Approach to Network Configuration Verification Batfish, 2017 Sem 3 Isolation Firecracker NSDI 2020 ; gVisor design docs Sem 2 ML systems Megatron-LM; ZeRO; Efficient Memory Management for LLM Serving with PagedAttention vLLM, 2023 Sem 3 Agent security Not what you've signed up for: indirect prompt injection Greshake et al., 2023 ; AgentDojo 2024 ; Defeating Prompt Injections by Design CaMeL, 2025 Sem 3 Research skills to practice spread over all four semesters Finding a gap: for each paper, write what it did not evaluate. Writing: follow the structure of a 6-page workshop paper problem, motivation with numbers, design, evaluation, related work . Figures: CDFs and time series with clearly labeled axes; reuse your benchmarking toolkit. Getting feedback: email authors with specific questions, post drafts publicly, ask for reviews in research communities. Weekly schedule fits your three-block day The program needs about 15 hours a week, taken from your 6 PM – 3 AM personal block: 9 hours on weekdays and 6 on weekends. Times are flexible; the weekly hour count is what matters. Day Time IST Hours Activity Monday 9 – 11 PM 2 Lectures and reading for the current course Tuesday 9 – 11 PM 2 Lab / project coding Wednesday 9 – 10 PM 1 Paper of the week first and second pass Thursday 9 – 11 PM 2 Lab / project coding Friday — 0 Rest; buffer for work overruns Saturday 11 AM – 2 PM 3 Deep work: hardest lab or thesis experiments Sunday 11 AM – 2 PM, 9 – 11 PM 3 + 2 Deep work; then paper summary + weekly review Adjustments In release crunches at work, drop to the minimum: Wednesday paper + one weekend session ~4 hours . Never skip two weeks in a row. From month 13, Saturday becomes a thesis day. Every 6 weeks, take one full week off. That keeps two years sustainable. Weekly review 15 minutes, Sunday night What did I finish this week? What is blocked, and why? Is the current course on track for its 6-month window? Next week's single most important task Assessment, milestones, and credibility without a degree Without a university, your proof is public work. Aim for evidence that a hiring manager or PhD advisor can check in five minutes. How to grade yourself per course Component Weight Pass bar Labs completed 40% All official labs pass their test suites Course project 40% Working repo + 2–4 page report with measured results Paper summaries 10% One per week, no gaps longer than 2 weeks Public write-up 10% One blog post explaining what you learned End-of-program portfolio checklist 10 course project repos with READMEs and reports 1 thesis 40–60 pages and its code, public on GitHub 1 arXiv preprint or workshop paper submission from the thesis At least 1 merged patch in a real project FRRouting, containerlab, Batfish, gVisor, or similar Self-Directed Research Program: MS by Research and PhD Track Sep 30, 2026 · @danisherror Program overview This program trains you to produce publishable research, not to finish courses. It has two tracks: an MS by Research equivalent about 2.5 years, 1–2 papers and a thesis and a PhD equivalent about 5 years total, 3–5 papers and a dissertation . The PhD track continues from the MS track; you decide at the MS gate whether to go on. How it differs from the M.Tech plan M.Tech plan This research program Main output Skills + projects New knowledge: peer-reviewed papers Coursework ~70% of time ~20%, front-loaded in year 1 Success measure Labs pass, projects work Reviewers at SIGCOMM/NSDI/USENIX accept the work Time budget ~15 h/week ~18–20 h/week research needs long uninterrupted blocks Research area:Dependable infrastructure for AI — networks and agent-execution systems that fail safely. It combines your control-plane depth TI-LFA, IS-IS-SR with your agent-security work approval-gated MCP, sandboxing . An honest limit. Self-study cannot grant a degree, and research without feedback drifts. This program therefore builds in external reviewers, collaborators, and real paper submissions from year 1, and ends with a path to register the work for a formal degree last section . Both tracks share the first 30 months; at Gate 2 you either stop with a completed MS thesis or continue into the PhD phases. Phase 1 months 1–10 : Research foundations Phase 1 replaces the coursework year of a research degree: four research-skill modules, a breadth survey, and a qualifying exam you set for yourself and have someone else grade. R1. Reading and critiquing research months 1–3 Method: S. Keshav's three-pass reading; for each paper, write a review in conference format summary, strengths, weaknesses, questions to authors, accept/reject . Volume: 3 papers/week for 12 weeks ~36 reviews . Calibration: pick papers from venues that publish reviews or public discussion e.g., OpenReview and compare your review with the real ones. Output: a review notebook on GitHub. R2. Experimental methodology months 2–5 Topics Choosing baselines that a reviewer will accept Workloads: real traces vs synthetic; where to find public datacenter and training-job traces Statistics: confidence intervals, repeated runs, variance sources on shared machines Simulation vs emulation vs testbed: ns-3, containerlab/Mininet, real hardware; validating a simulator against reality Reproducibility: artifact packaging, scripts that regenerate every figure Project: reproduce one published result for example, a DCQCN or ECMP-imbalance figure in simulation and write up where your numbers differ and why. R3. Research writing months 4–10, continuous Study: Simon Peyton Jones's talk How to Write a Great Research Paper; Jim Kurose's advice on writing paper introductions. Practice: rewrite the introduction of your reproduction report five times, getting feedback each time. Learn LaTeX and the ACM/USENIX templates; keep a figure-generation pipeline in Python. R4. Breadth survey months 5–9 Write a 15–20 page survey: "Failure handling in AI infrastructure: networks, collectives, and agent execution". Cover ~60–80 papers, organize them into a taxonomy, and end with 5–10 open problems. This becomes your map for the next four years and can be posted on arXiv. Qualifying exam equivalent month 10 Part Format Who grades it Written Answer 3 of 5 questions on your survey area, 48 hours, open book An external mentor see the advisors section Oral 45-minute talk on your survey + open problems, then 30 minutes of questions Mentor + 1–2 engineers or researchers Pass bar You can defend why each open problem matters and how you would evaluate a solution Mentor's judgment If you fail, spend 6 more weeks on the weak area and retake it. Do not start the main research agenda before passing. Research agenda Central question: How do we build AI infrastructure whose failures, whether hardware faults in the network or wrong actions by an AI agent, are contained quickly and provably, without a human debugging every incident? A PhD is a chain of papers that each answer one part of this question. The first two papers form the MS thesis; papers 3–5 extend it into a PhD. Paper working title Research question Method Target venue Track P1 Measuring the cost of link failures in AI training fabrics How much job time does one link or switch failure cost at different scales, and where does it go detection, reconvergence, stragglers ? Trace-driven simulation of rail-optimized Clos fabrics; validation on a small rented GPU cluster HotNets or APNet workshop MS P2 Collective-aware fast reroute for GPU fabrics Can precomputed, collective-aware backup paths TI-LFA-style cut the time lost to failures? Can every agent action be tied to an explicit, revalidated user approval, even under prompt injection? TLA+ protocol spec, implementation from your MCP work, adversarial evaluation USENIX Security or CCS PhD P5 Unified containment for AI infrastructure Can one framework cover both network faults and agent faults detect, contain, recover, audit ? Synthesis of P2–P4 into one system; end-to-end evaluation OSDI or NSDI PhD Why this chain works: P1 produces the numbers that motivate P2. P3 and P4 reuse the approval and verification ideas from each other. P5 is the dissertation's unifying contribution. Each paper can stand alone if a later one fails. Keep a backup topic. If P1 shows that failure cost is already small in practice, pivot to P3/P4 agent safety as the main line; the survey from Phase 1 will show other open problems. MS by Research track months 11–30 The MS track ends with a thesis built on P1 and P2, at least one submitted paper, and a public defense. Stages Months 11–12: Research proposal 5–8 pages . Problem, why it matters with numbers from your survey , related work, planned approach, evaluation plan, risks. Your mentor must approve it. Months 13–18: P1 measurement study. Build the simulator, validate it on a small real setup, publish the tool as open source. Submit P1 to a workshop. Months 19–27: P2 system. Design the fast-reroute scheme, prototype it, evaluate at scale. Submit P2 to a main conference expect rejection on the first try; revise and resubmit . Months 28–30: Thesis writing and defense. Write the thesis and give a 45-minute public defense online meetup, recorded with your mentor and two outside reviewers asking questions. Background: AI training traffic, datacenter fabrics, fast reroute LFA, RLFA, TI-LFA Measurement study P1 Collective-aware fast reroute: design P2 Implementation Evaluation Related work Limitations and future work Appendix: artifact and reproduction instructions MS exit criteria Proposal approved by mentor P1 submitted workshop and posted on arXiv P2 submitted to a main venue at least once Code and data public, with a script that regenerates every figure Thesis written and defended Decision at month 30: continue to the PhD track only if you enjoyed the uncertainty of research, P2 got substantive reviews even if rejected , and you still have a question you want to answer. Otherwise, the MS is a strong finish on its own. PhD track months 31–60 The PhD track adds P3, P4, and P5 on top of the MS work and ends with a dissertation whose central claim ties them together. Stages Months 31–33: Dissertation proposal 15–20 pages . Thesis statement, completed work P1, P2 , planned work P3–P5 , timeline, and risks. Defend it orally before a committee of three: your mentor plus two outside researchers. Months 34–42: P3, verified AI-generated network change. Build the pipeline and a public benchmark of change requests; submit to NSDI or CoNEXT. Months 40–50: P4, provable agent authorization. Formal spec in TLA+, implementation, adversarial evaluation; submit to USENIX Security or CCS. Overlaps P3 by design so that a rejection cycle on one keeps the other moving. Months 50–56: P5, unified containment. The capstone system. Months 56–60: Dissertation and defense. Dissertation structure 120–180 pages Introduction and thesis statement, e.g.: "Failures in AI infrastructure, whether in the network or in agent actions, can be contained within bounded time and with verifiable authorization by combining precomputed local recovery, pre-deployment verification, and explicit approval protocols." Background and survey from R4, updated Part I: Network failures P1, P2 Part II: Agent and automation failures P3, P4 Part III: Unified containment P5 Related work, limitations, future directions Publication targets Level Count by month 60 Minimum 3 papers accepted, at least 1 at a top venue SIGCOMM, NSDI, OSDI, SOSP, USENIX Security, IEEE S&P, CCS Strong 4–5 accepted, 2+ at top venues, one artifact badge e.g., artifact evaluation at USENIX Also expected Serve as a reviewer or shadow PC member at least once; 2+ public talks Rejections are normal: many top-venue papers are rejected once or twice before acceptance. Budget for one resubmission per paper in the timeline. Advisors, collaborators, compute, and testbeds The single biggest risk of independent research is working alone. Secure an external mentor by month 6 and a co-author by month 18. Finding a mentor the advisor substitute From your survey, list 15–20 researchers faculty at IISc, IITs, or abroad; researchers at Microsoft Research India, Google, Meta, NVIDIA whose papers you cite most. Email them with something concrete: a reproduction result that differs from their paper, a bug in their artifact, or a short question about their evaluation. Never send a generic "please mentor me" email. Ask for a light commitment: a 30-minute call once a month. Expect a low response rate; one yes out of twenty is a good outcome. Other sources of feedback Industry colleagues at HPE with networking depth: ask them to review drafts and serve on your "committee". Research communities and workshops: submit to workshops and student/poster sessions many systems conferences have them for early feedback. Open-source maintainers: FRRouting, containerlab, and Batfish developers can review the practical side of your work. Apply to research programs that take working professionals as collaborators or visiting researchers; check what Microsoft Research India and IISc/IIT labs currently offer. Compute and testbeds Need Option Notes Network emulation containerlab, Mininet, BMv2 on your Mac or a Linux VM Free; enough for P2, P3 prototypes Large-scale simulation ns-3 or a custom flow-level simulator on a rented multi-core cloud VM Pay per hour; keep runs scripted Real GPU traffic traces A few hours on rented 4–8 GPU cloud instances Budget carefully; capture traces once and reuse them Research testbeds CloudLab, FABRIC, Chameleon US research testbeds Access usually requires an academic affiliation; a collaborator can give you access Cloud credits Research credit programs from cloud providers Usually need a proposal; your research proposal doubles as the application Budget: plan roughly for cloud spending of a few thousand rupees a month, rising during evaluation phases. Collaborating with an academic lab is the best way to cut this. Weekly routine and year-end checkpoints Research needs long, uninterrupted blocks, so this routine puts most hours on weekends and keeps weekdays for reading, writing, and small tasks. Total: about 18–20 hours/week. Day Time IST Hours Activity Monday 9 – 11 PM 2 Read: 1 paper, write its review Tuesday 9 – 11 PM 2 Code: small experiment changes, bug fixes Wednesday 9 – 10:30 PM 1.5 Write: 500 words on the current paper or thesis Thursday 9 – 11 PM 2 Code or analyze last weekend's results Friday — 0 Rest Saturday 10 AM – 3 PM 5 Deep work: design, large experiments Sunday 10 AM – 2 PM, 9 – 11 PM 4 + 2 Deep work; then research log + plan next week Habits that matter more than hours Keep a dated research log: every experiment, its hypothesis, the result, and what you concluded. This becomes your thesis's raw material. Start long simulations on Sunday night so they finish during the week. Monthly: a 30-minute mentor call with a 1-page update sent two days before. Before each paper deadline, take 1–2 weeks of leave from work if possible; the last weeks before submission are the most intense. Year-end checkpoints End of year 1: qualifier passed, survey on arXiv, mentor secured, research proposal approved End of year 2: P1 submitted, P2 prototype working with first results End of year 2.5: MS thesis defended; continue-or-stop decision made End of year 3: dissertation proposal defended, P3 submitted End of year 4: P4 submitted, at least 2 papers accepted overall End of year 5: P5 done, dissertation defended Evaluating progress and turning it into a formal degree Peer review is the real grade: a paper accepted at a top venue is the same evidence whether or not you are enrolled anywhere. Track these signals every 6 months. Signal Healthy Warning Research log Entries most weeks Gaps of a month or more Submissions One submission every 6–9 months after year 1 Nothing submitted in 12 months Reviews received Reviewers engage with the idea, criticize the evaluation Reviewers say the problem is not important Mentor contact Monthly calls happening No contact in 3 months Your motivation You think about the problem outside scheduled hours Only deadlines make you work If reviewers repeatedly say the problem is not important, revisit the agenda with your mentor instead of polishing the same paper. Paths to a formal degree later Your published work makes these routes much stronger. Rules and eligibility change, so check each institute's current admission pages. Part-time or external PhD at IITs/IISc for working professionals: some programs let employees register while working, usually with a short coursework requirement and sometimes an employer sponsorship letter. Your papers and a willing faculty co-author make admission far more likely. PhD by publication: some universities mainly in the UK and parts of Europe and Australia award a PhD based on a portfolio of published papers plus a written commentary; check eligibility for external candidates. Industry research roles: papers at NSDI/SIGCOMM/USENIX Security can qualify you for research-engineer roles at industry labs even without a PhD. Full-time admission later: if your situation changes, a strong publication record also helps for a funded full-time PhD in India or abroad. The mentor you find in year 1 matters here too: a faculty member who has co-authored with you is the most natural PhD advisor if you register formally.