Got it, this is your self-directed M.Tech in Networked Systems & Security for AI Infrastructure. That changes things from my earlier generic answer, so here's a schedule built specifically around this program's Semester 1 (C1 OS, C2 Advanced Networks, C3 Math, C4 Go), fitted to a 5-hour workday.
How much time to use
Your plan is designed for about 15 hours a week. With a 5-hour job you have room for more, but I'd keep 15 hours as the committed load and treat the extra capacity as a buffer (Friday catch-up, longer lab sessions when something breaks). Two years is long, and the plan's own rule of a week off every 6 weeks only works if the weekly load stays sustainable.
The weekly routine (same slots your doc uses; shift them if your work hours differ)
Day
Time
Hrs
What
Mon
9β11 PM
2
Main course: lectures and reading
Tue
9β11 PM
2
Main course: lab coding
Wed
9β10 PM
1
Paper of the week (first and second pass)
Thu
9β11 PM
2
Side course: lectures, exercises
Fri
9β11 PM
0β2
Buffer only: catch up if behind, otherwise rest
Sat
11 AMβ2 PM
3
Deep work: main course lab or project
Sun
11 AMβ2 PM
3
Side course project or practice
Sun
9β11 PM
2
Paper summary + 15-min weekly review
The key idea is running two tracks at once instead of four courses in parallel: one heavy main course (OS, then Networks) and one lighter side course (Go, then Math). Two-hour sessions split across four courses would mean constant context switching and no lab ever getting finished.
Semester 1 week-by-week (26 weeks, with rest weeks at 7, 14, and 21) Weeks
Main track
Side track
1β6
C1 OS: OSTEP reading + xv6 labs (util + syscall, page tables, traps, COW fork, one lab per week)
Wrap up C3 notes and write the Semester 1 blog post
A few reasons behind the ordering: Go comes first because it's short and you'll need it for the MIT 6.5840 labs in Semester 2. The benchmarking toolkit (C3) is finished by week 16 so it's ready for the C2 congestion-control study. OS comes before Networks because the xv6 networking lab and container work make the TCP stack labs easier to reason about. And your networking background means C2 can move faster than its topic list suggests.
Papers for this semester
Start week 1 with Keshav's "How to Read a Paper," then the three datacenter papers the plan assigns to Semester 1 (Al-Fares fat-tree, VL2, Jupiter Rising) in weeks 2β4. After that, pick papers tied to whatever the main track is covering, such as a classic OS or virtualization paper during C1 and a TCP or datacenter congestion paper during C2.
When work gets busy
Follow the minimum from your plan: Wednesday paper plus one weekend session (about 4 hours), and never skip two weeks in a row. If a course falls more than a couple of weeks behind this table, drop the lowest-value project part and keep the labs, then use the next rest week as a catch-up week instead of a full break.
Want me to add this as a "Semester 1 schedule" tab in your program doc so it sits alongside the course details?
Self-Directed M.Tech: Networked Systems & Security for AI Infrastructure
Sep 30, 2026 Β· @danisherror
Program overview
This is a 2-year, 4-semester self-study program at roughly 15 hours/week (~1,500 hours total), equivalent in depth to an M.Tech in Computer Systems with a research thesis. The goal is to become someone who can build and secure the infrastructure AI runs on: GPU datacenter networks, sandboxed agent execution, and AI-assisted network automation.
Why this track fits you. It builds on what you already do well (C, control planes, IS-IS/SR, TI-LFA/FRR) instead of restarting in crowded pure ML. It connects directly to your MCP approval system, dev-VM sandbox, Process Network Inspector, and the Apstra work at HPE.
How it maps to a real M.Tech.
Real M.Tech component
Equivalent here
Core courses (~6) 6 core courses across Semesters 1β3
Electives (~3) 3 electives in Semester 3, chosen from a list
Lab / assignments
1β2 graded projects per course
Seminar
Paper-reading practice: 1 paper/week, written summaries
Thesis (~1 year) Research project from Semester 3, full-time focus in Semester 4
Degree certificate
Public portfolio: GitHub repos, technical blog, open-source commits, a workshop paper or arXiv preprint Rules that keep it rigorous.
Every course ends with a project that has a written report (2β4 pages) and a public repo.
Use free university courses (MIT, CMU, Stanford, Berkeley) as the lecture backbone; do their labs, not just the videos. Write in C, Go, or Python depending on the layer; C and Go carry most of the systems work.
Course URLs and editions change; search the course code to find the current offering.
Program structure
[embedded content: program roadmap Β· 4 semesters, 3 gates] Semesters 1 and 2 build the systems base; the thesis starts at month 13 beside the AI-layer courses and takes over fully in Semester 4. Do not pass a gate until its condition is met.
Semester 1 (months 1β6): Systems foundations Semester 1 closes the gaps under your networking strength: OS internals, modern networking beyond routing, the math research needs, and Go.
C1. Operating Systems & Systems Programming
Backbone: MIT 6.1810 (formerly 6.S081, xv6 labs); book Operating Systems: Three Easy Pieces (free online). Topics
Processes, threads, scheduling, context switches
Virtual memory, page tables, TLBs, copy-on-write
System calls, traps, interrupts, the kernel/user boundary
Mini-container runtime in C/Go: build a tool that uses namespaces, cgroups v2, pivot_root, seccomp, and dropped capabilities to run a process. Write a threat model listing what it does and does not isolate. This formalizes your recent container investigation.
C2. Advanced Computer Networks
Backbone: Stanford CS144 (build a TCP stack); Computer Networks: A Systems Approach by Peterson & Davie (free online); Stanford CS244 or Princeton COS 561 reading lists for advanced papers.
Topics
TCP internals: congestion control (Reno, CUBIC, BBR), flow control, retransmission
Network measurement and telemetry: sFlow, INT, gNMI streaming
Congestion in datacenters: ECN, DCTCP, incast
Routing at scale: BGP in the datacenter (RFC 7938), EVPN-VXLAN Projects
Complete the CS144 TCP implementation labs.
Congestion-control lab: in Mininet or ns-3, compare CUBIC, BBR, and DCTCP under incast and report the results with graphs.
EVPN-VXLAN fabric in containerlab: build a 2-spine/4-leaf fabric with FRRouting, then break links and measure convergence.
C3. Mathematics for Systems Research
Backbone: MIT 6.041 (Probability), Gilbert Strang's MIT 18.06 (Linear Algebra), selected queueing theory chapters (Mor Harchol-Balter, Performance Modeling and Design of Computer Systems).
Queueing theory: M/M/1, Little's law, load vs latency
Linear algebra essentials for ML (matrices, gradients, tensor shapes)
Graph algorithms: shortest paths, max-flow, spanning trees (you know SPF; add flow and cuts) Statistics for experiments: confidence intervals, variance, how to compare two systems fairly
Project
Statistical benchmarking toolkit (Python): run a workload N times, compute confidence intervals, and draw CDF plots. You will reuse it in every later experiment.
C4. Go for Systems (short course, ~6 weeks)
Backbone:The Go Programming Language (Donovan & Kernighan); Go by Example; Go's official concurrency material.
Topics: goroutines, channels, context cancellation, gRPC/protobuf, profiling with pprof, writing CLIs and daemons.
Project
gNMI telemetry collector in Go: subscribe to streaming telemetry from a containerlab device (SR Linux or cEOS), store it in PostgreSQL or a time-series DB, and alert on link flaps. This maps directly to Apstra device-agent work.
Semester 2 (months 7β12): Core specialization Semester 2 builds the three pillars of the specialization: distributed systems, networks for AI clusters, and systems security.
C5. Distributed Systems
Backbone: MIT 6.5840 (formerly 6.824) with its Go labs; Designing Data-Intensive Applications (Kleppmann). Replicated approval store: rebuild the approval/audit store from your MCP dev-VM project on top of your Raft implementation, so approvals survive node failure and the hash-chained audit log stays consistent.
C6. Datacenter Networking for AI Clusters
This is the course that makes you distinct. No single university course covers it fully, so the backbone is papers plus labs.
Backbone: papers listed under the reading list (Jupiter, fat-tree, DCQCN, RDMA at scale, HPCC, Meta and Alibaba AI-training networks); NVIDIA NCCL documentation; Ultra Ethernet Consortium public material.
Topics
How distributed training works: data/tensor/pipeline parallelism, all-reduce, all-to-all
Collective communication: ring and tree all-reduce, NCCL internals
Failure handling: link flaps, stragglers, how one failure stalls a training job
Fast reroute in the fabric: applying LFA/TI-LFA thinking to AI fabrics
Telemetry and root-cause analysis for training-job slowdowns
Projects
All-reduce simulator (Python or Go): simulate ring vs tree all-reduce on a fat-tree in ns-3 or a custom flow-level simulator; measure job completion time under one failed link. ECMP collision study: show how a few large flows collide under hash-based ECMP and evaluate a flowlet- or spraying-based fix.
Write-up: a 3-page survey "Failure recovery in AI training fabrics" β this often becomes the seed of the thesis.
C7. Systems Security
Backbone: Stanford CS155 or Berkeley CS161; MIT 6.5660 (formerly 6.858) labs; Ross Anderson, Security Engineering (3rd ed.). Complete the MIT 6.5660 lab series (privilege separation, web security).
Sandbox comparison study: run the same untrusted workload under Docker, gVisor, Firecracker, and Kata on Linux; compare isolation guarantees, startup time, and syscall overhead with your benchmarking toolkit.
Formal threat model for your MCP CRUD system: attack tree, trust boundaries, and a test suite of malicious-agent scenarios.
Semester 3 (months 13β18): AI systems, programmable networks, electives
Semester 3 adds the AI layer on top of your systems base and starts the thesis in parallel (about 5 of the 15 weekly hours go to research from month 13).
C8. ML Systems (how AI actually runs) Backbone: Andrej Karpathy's Neural Networks: Zero to Hero (build GPT from scratch); CMU 10-414/714 Deep Learning Systems (build a mini PyTorch); Stanford CS336 Language Modeling from Scratch.
Topics
Neural network basics, backpropagation, transformers and attention
Tensors, autograd, GPU memory hierarchy, kernels
Training at scale: data/tensor/pipeline parallelism, ZeRO, checkpointing
Build a small GPT from scratch in PyTorch and train it on a small dataset (your M2 Mac handles small models; use a cloud GPU for a few hours when needed).
Network-trace generator from training: instrument a 2β4 GPU data-parallel job (rented) and capture its all-reduce traffic pattern; feed it into your C6 simulator.
C9. Programmable Networks & Fast Packet Processing
Backbone: p4lang/tutorials on GitHub; Learning eBPF (Liz Rice); DPDK documentation; George Varghese, Network Algorithmics.
Complete the P4 tutorials; implement a P4 program that does per-flow telemetry.
eBPF process-network observer on Linux: a Linux counterpart to your Process Network Inspector that maps sockets and flows to PIDs via eBPF, strictly view-only.
Fast-reroute in P4: implement data-plane link-failure detection and local reroute in BMv2, and compare failover time against control-plane recomputation.
C10. Electives (pick 3) Elective
Backbone
Project
Best if thesis is on
E1. Security of LLM agents
OWASP Top 10 for LLM Applications; papers on prompt injection and CaMeL; AgentDojo benchmark
Red-team your MCP CRUD system with injected tool outputs; measure how many attacks the approval layer blocks
Agent security
E2. Network verification
Batfish docs; papers on header space analysis and config verification
Use Batfish to verify reachability and ECMP properties of your EVPN fabric before and after a change Safe AI network automation
E3. Formal methods (TLA+) Leslie Lamport's TLA+ video course; Specifying Systems
Write a TLA+ spec of your proposalβapprovalβrevalidationβexecute protocol and model-check race conditions
Agent security or safe automation
E4. Intent-based networking
Apstra public docs; RFC 9315 (intent-based networking concepts); SONiC docs Build a tiny intent engine: YAML intent β rendered FRR config β Batfish check β deploy to containerlab
Safe AI network automation
E5. Performance engineering
Brendan Gregg's books; Systems Performance
Profile and speed up a real open-source network daemon (e.g., FRR isisd) and upstream a patch
AI fabric reliability
E6. Storage systems
CMU 15-445 (databases); papers on distributed file systems Build a checkpoint store for training jobs and measure recovery time
AI fabric reliability
Recommended default set: E1 + E2 + E3. They cover security, verification, and formal reasoning, which strengthen any of the thesis topics below.
Research projects and thesis options
Pick one thesis topic by month 12 and work on it from month 13 to month 24. The first option is the strongest match for your background; the other three are solid alternatives.
Thesis option A (recommended): Fast failure recovery for AI training fabrics
Research question: When a link or switch fails in a GPU cluster fabric, how much training time is lost, and can local fast-reroute (in the spirit of TI-LFA) plus collective-aware traffic engineering cut that loss significantly?
Why you: almost nobody working on AI networking has shipped TI-LFA in production code. You have.
Method
Build a flow-level simulator of a rail-optimized Clos fabric driven by real all-reduce traces (from C8).
Measure baseline: ECMP rehash plus control-plane reconvergence after failures.
Design a scheme: precomputed backup paths per collective group, placed so the backup avoids congested links.
Implement a prototype in P4 (BMv2) or FRR plus containerlab for the control-plane part.
Evaluate: job completion time, tail latency of collectives, backup-path computation cost, at 64 to 4,096 simulated GPUs.
Target venues: HotNets, the SIGCOMM or NSDI workshops, APNet; arXiv preprint regardless.
Research question: Can an LLM agent propose network changes that are provably safe before deployment, using verification (Batfish), formal intent, and human approval gates?
Method: build a pipeline of natural-language intent β LLM-generated config β Batfish verification β diff-based human approval β staged rollout with automatic rollback. Create a benchmark of 100+ change requests including deliberately risky ones, and measure how many unsafe changes each safety layer catches.
Why you: combines Apstra-style intent networking with your approval-gated MCP design. Directly useful at HPE.
Thesis option C: Verifiable authorization for AI agents acting on real systems
Research question: How can a system prove that every action an agent took was explicitly authorized by a user, even when the agent is manipulated by prompt injection?
Method: formalize your proposal β approval β revalidation β execute protocol in TLA+; implement it with canonical proposal hashes and a transparency-log-style audit; attack it with injected tool outputs (AgentDojo-style); measure attack success rate and approval overhead.
Why you: this is your MCP CRUD and dev-VM work turned into research.
Thesis option D: Sandboxing untrusted agent code
Research question: What is the right isolation boundary (container, gVisor, microVM) for agent-executed code, and what capability model lets agents do useful work without broad access?
Method: systematically compare escape surface, syscall coverage, and startup latency; design an effect-declaration system (from your dev-VM project) and enforce it with seccomp or eBPF LSM.
Smaller research projects (portfolio pieces)
Complete at least two of these alongside the courses; each is 4β8 weeks.
Project
Output
Links to
Measure ECMP imbalance under AI-style elephant flows
Blog post + repo
C6, thesis A
Upstream a fix or feature to FRRouting isisd (e.g., TI-LFA improvements)
Merged patch
Your IS-IS-SR expertise LLM agent that diagnoses a broken containerlab fabric from telemetry
Demo + write-up
C4, C8, thesis B
Prompt-injection test suite for MCP servers
Open-source tool
E1, thesis C
TLA+ spec of an approval protocol, with a found bug documented
Blog post
E3, thesis C
Paper reading list and research skills
Read one paper a week (~100 over two years) and write a 1-page summary for each: problem, key idea, evaluation, weakness, and one follow-up idea. Start with S. Keshav's short guide How to Read a Paper (three-pass method).
Where the field publishes: SIGCOMM, NSDI, HotNets, CoNEXT (networking); OSDI, SOSP, EuroSys, USENIX ATC (systems); USENIX Security, IEEE S&P, CCS (security); MLSys (ML systems). Area
Papers to start with
Read in
Datacenter networks
A Scalable, Commodity Data Center Network Architecture (Al-Fares et al., 2008); VL2 (2009); Jupiter Rising (Google, 2015) Sem 1
Lossless / RDMA networking
Congestion Control for Large-Scale RDMA Deployments (DCQCN, 2015); RDMA over Commodity Ethernet at Scale (2016); HPCC (2019); Swift (2020) Sem 2
AI training networks
RDMA over Ethernet for Distributed AI Training at Meta Scale (2024); Alibaba HPN (2024) Sem 2
Distributed systems
Raft (2014); Paxos Made Simple; MapReduce; GFS; Spanner; Borg Header Space Analysis (2012); A General Approach to Network Configuration Verification (Batfish, 2017)
Sem 3
Isolation
Firecracker (NSDI 2020); gVisor design docs Sem 2
ML systems
Megatron-LM; ZeRO; Efficient Memory Management for LLM Serving with PagedAttention (vLLM, 2023) Sem 3
Agent security
Not what you've signed up for: indirect prompt injection (Greshake et al., 2023); AgentDojo (2024); Defeating Prompt Injections by Design (CaMeL, 2025) Sem 3
Research skills to practice (spread over all four semesters)
Finding a gap: for each paper, write what it did not evaluate.
Writing: follow the structure of a 6-page workshop paper (problem, motivation with numbers, design, evaluation, related work).
Figures: CDFs and time series with clearly labeled axes; reuse your benchmarking toolkit.
Getting feedback: email authors with specific questions, post drafts publicly, ask for reviews in research communities.
Weekly schedule (fits your three-block day) The program needs about 15 hours a week, taken from your 6 PM β 3 AM personal block: 9 hours on weekdays and 6 on weekends. Times are flexible; the weekly hour count is what matters.
Day
Time (IST) Hours
Activity
Monday
9 β 11 PM
2
Lectures and reading for the current course
Tuesday
9 β 11 PM
2
Lab / project coding
Wednesday
9 β 10 PM
1
Paper of the week (first and second pass)
Thursday
9 β 11 PM
2
Lab / project coding
Friday
β
0
Rest; buffer for work overruns
Saturday
11 AM β 2 PM
3
Deep work: hardest lab or thesis experiments
Sunday
11 AM β 2 PM, 9 β 11 PM
3 + 2
Deep work; then paper summary + weekly review
Adjustments
In release crunches at work, drop to the minimum: Wednesday paper + one weekend session (~4 hours). Never skip two weeks in a row.
From month 13, Saturday becomes a thesis day. Every 6 weeks, take one full week off. That keeps two years sustainable.
Weekly review (15 minutes, Sunday night)
What did I finish this week?
What is blocked, and why?
Is the current course on track for its 6-month window?
Next week's single most important task
Assessment, milestones, and credibility without a degree
Without a university, your proof is public work. Aim for evidence that a hiring manager or PhD advisor can check in five minutes.
How to grade yourself (per course) Component
Weight
Pass bar
Labs completed
40%
All official labs pass their test suites
Course project
40%
Working repo + 2β4 page report with measured results
Paper summaries
10%
One per week, no gaps longer than 2 weeks
Public write-up 10%
One blog post explaining what you learned
End-of-program portfolio checklist 10 course project repos with READMEs and reports
1 thesis (40β60 pages) and its code, public on GitHub
1 arXiv preprint or workshop paper submission from the thesis
At least 1 merged patch in a real project (FRRouting, containerlab, Batfish, gVisor, or similar)
Self-Directed Research Program: MS by Research and PhD Track
Sep 30, 2026 Β· @danisherror
Program overview
This program trains you to produce publishable research, not to finish courses. It has two tracks: an MS by Research equivalent (about 2.5 years, 1β2 papers and a thesis) and a PhD equivalent (about 5 years total, 3β5 papers and a dissertation). The PhD track continues from the MS track; you decide at the MS gate whether to go on.
How it differs from the M.Tech plan
M.Tech plan
This research program
Main output
Skills + projects
New knowledge: peer-reviewed papers Coursework
~70% of time
~20%, front-loaded in year 1
Success measure
Labs pass, projects work
Reviewers at SIGCOMM/NSDI/USENIX accept the work
Time budget
~15 h/week
~18β20 h/week (research needs long uninterrupted blocks)
Research area:Dependable infrastructure for AI β networks and agent-execution systems that fail safely. It combines your control-plane depth (TI-LFA, IS-IS-SR) with your agent-security work (approval-gated MCP, sandboxing).
An honest limit. Self-study cannot grant a degree, and research without feedback drifts. This program therefore builds in external reviewers, collaborators, and real paper submissions from year 1, and ends with a path to register the work for a formal degree (last section).
Both tracks share the first 30 months; at Gate 2 you either stop with a completed MS thesis or continue into the PhD phases.
Phase 1 (months 1β10): Research foundations Phase 1 replaces the coursework year of a research degree: four research-skill modules, a breadth survey, and a qualifying exam you set for yourself and have someone else grade.
R1. Reading and critiquing research (months 1β3)
Method: S. Keshav's three-pass reading; for each paper, write a review in conference format (summary, strengths, weaknesses, questions to authors, accept/reject).
Volume: 3 papers/week for 12 weeks (~36 reviews). Calibration: pick papers from venues that publish reviews or public discussion (e.g., OpenReview) and compare your review with the real ones.
Output: a review notebook on GitHub.
R2. Experimental methodology (months 2β5)
Topics
Choosing baselines that a reviewer will accept
Workloads: real traces vs synthetic; where to find public datacenter and training-job traces
Statistics: confidence intervals, repeated runs, variance sources on shared machines
Simulation vs emulation vs testbed: ns-3, containerlab/Mininet, real hardware; validating a simulator against reality
Reproducibility: artifact packaging, scripts that regenerate every figure
Project: reproduce one published result (for example, a DCQCN or ECMP-imbalance figure) in simulation and write up where your numbers differ and why.
R3. Research writing (months 4β10, continuous)
Study: Simon Peyton Jones's talk How to Write a Great Research Paper; Jim Kurose's advice on writing paper introductions.
Practice: rewrite the introduction of your reproduction report five times, getting feedback each time.
Learn LaTeX and the ACM/USENIX templates; keep a figure-generation pipeline in Python.
R4. Breadth survey (months 5β9) Write a 15β20 page survey: "Failure handling in AI infrastructure: networks, collectives, and agent execution". Cover ~60β80 papers, organize them into a taxonomy, and end with 5β10 open problems. This becomes your map for the next four years and can be posted on arXiv.
Qualifying exam equivalent (month 10) Part
Format
Who grades it
Written
Answer 3 of 5 questions on your survey area, 48 hours, open book
An external mentor (see the advisors section)
Oral
45-minute talk on your survey + open problems, then 30 minutes of questions
Mentor + 1β2 engineers or researchers
Pass bar
You can defend why each open problem matters and how you would evaluate a solution
Mentor's judgment
If you fail, spend 6 more weeks on the weak area and retake it. Do not start the main research agenda before passing. Research agenda
Central question: How do we build AI infrastructure whose failures, whether hardware faults in the network or wrong actions by an AI agent, are contained quickly and provably, without a human debugging every incident?
A PhD is a chain of papers that each answer one part of this question. The first two papers form the MS thesis; papers 3β5 extend it into a PhD.
#
Paper (working title) Research question
Method
Target venue
Track
P1
Measuring the cost of link failures in AI training fabrics
How much job time does one link or switch failure cost at different scales, and where does it go (detection, reconvergence, stragglers)?
Trace-driven simulation of rail-optimized Clos fabrics; validation on a small rented GPU cluster
HotNets or APNet (workshop) MS
P2
Collective-aware fast reroute for GPU fabrics
Can precomputed, collective-aware backup paths (TI-LFA-style) cut the time lost to failures? Can every agent action be tied to an explicit, revalidated user approval, even under prompt injection?
TLA+ protocol spec, implementation from your MCP work, adversarial evaluation
USENIX Security or CCS
PhD
P5
Unified containment for AI infrastructure
Can one framework cover both network faults and agent faults (detect, contain, recover, audit)?
Synthesis of P2βP4 into one system; end-to-end evaluation OSDI or NSDI
PhD
Why this chain works: P1 produces the numbers that motivate P2. P3 and P4 reuse the approval and verification ideas from each other. P5 is the dissertation's unifying contribution. Each paper can stand alone if a later one fails.
Keep a backup topic. If P1 shows that failure cost is already small in practice, pivot to P3/P4 (agent safety) as the main line; the survey from Phase 1 will show other open problems.
MS by Research track (months 11β30) The MS track ends with a thesis built on P1 and P2, at least one submitted paper, and a public defense.
Stages
Months 11β12: Research proposal (5β8 pages). Problem, why it matters (with numbers from your survey), related work, planned approach, evaluation plan, risks. Your mentor must approve it.
Months 13β18: P1 measurement study. Build the simulator, validate it on a small real setup, publish the tool as open source. Submit P1 to a workshop.
Months 19β27: P2 system. Design the fast-reroute scheme, prototype it, evaluate at scale. Submit P2 to a main conference (expect rejection on the first try; revise and resubmit).
Months 28β30: Thesis writing and defense. Write the thesis and give a 45-minute public defense (online meetup, recorded) with your mentor and two outside reviewers asking questions.
Background: AI training traffic, datacenter fabrics, fast reroute (LFA, RLFA, TI-LFA)
Measurement study (P1)
Collective-aware fast reroute: design (P2)
Implementation
Evaluation
Related work
Limitations and future work
Appendix: artifact and reproduction instructions
MS exit criteria
Proposal approved by mentor
P1 submitted (workshop) and posted on arXiv
P2 submitted to a main venue at least once
Code and data public, with a script that regenerates every figure
Thesis written and defended
Decision at month 30: continue to the PhD track only if you enjoyed the uncertainty of research, P2 got substantive reviews (even if rejected), and you still have a question you want to answer. Otherwise, the MS is a strong finish on its own.
PhD track (months 31β60) The PhD track adds P3, P4, and P5 on top of the MS work and ends with a dissertation whose central claim ties them together.
Stages
Months 31β33: Dissertation proposal (15β20 pages). Thesis statement, completed work (P1, P2), planned work (P3βP5), timeline, and risks. Defend it orally before a committee of three: your mentor plus two outside researchers.
Months 34β42: P3, verified AI-generated network change. Build the pipeline and a public benchmark of change requests; submit to NSDI or CoNEXT.
Months 40β50: P4, provable agent authorization. Formal spec in TLA+, implementation, adversarial evaluation; submit to USENIX Security or CCS. Overlaps P3 by design so that a rejection cycle on one keeps the other moving.
Months 50β56: P5, unified containment. The capstone system.
Months 56β60: Dissertation and defense.
Dissertation structure (120β180 pages) Introduction and thesis statement, e.g.: "Failures in AI infrastructure, whether in the network or in agent actions, can be contained within bounded time and with verifiable authorization by combining precomputed local recovery, pre-deployment verification, and explicit approval protocols."
Background and survey (from R4, updated)
Part I: Network failures (P1, P2)
Part II: Agent and automation failures (P3, P4)
Part III: Unified containment (P5)
Related work, limitations, future directions
Publication targets
Level
Count by month 60
Minimum
3 papers accepted, at least 1 at a top venue (SIGCOMM, NSDI, OSDI, SOSP, USENIX Security, IEEE S&P, CCS)
Strong
4β5 accepted, 2+ at top venues, one artifact badge (e.g., artifact evaluation at USENIX)
Also expected
Serve as a reviewer or shadow PC member at least once; 2+ public talks
Rejections are normal: many top-venue papers are rejected once or twice before acceptance. Budget for one resubmission per paper in the timeline.
Advisors, collaborators, compute, and testbeds
The single biggest risk of independent research is working alone. Secure an external mentor by month 6 and a co-author by month 18.
Finding a mentor (the advisor substitute)
From your survey, list 15β20 researchers (faculty at IISc, IITs, or abroad; researchers at Microsoft Research India, Google, Meta, NVIDIA) whose papers you cite most. Email them with something concrete: a reproduction result that differs from their paper, a bug in their artifact, or a short question about their evaluation. Never send a generic "please mentor me" email.
Ask for a light commitment: a 30-minute call once a month.
Expect a low response rate; one yes out of twenty is a good outcome.
Other sources of feedback
Industry colleagues at HPE with networking depth: ask them to review drafts and serve on your "committee".
Research communities and workshops: submit to workshops and student/poster sessions (many systems conferences have them) for early feedback.
Open-source maintainers: FRRouting, containerlab, and Batfish developers can review the practical side of your work.
Apply to research programs that take working professionals as collaborators or visiting researchers; check what Microsoft Research India and IISc/IIT labs currently offer.
Compute and testbeds
Need
Option
Notes
Network emulation
containerlab, Mininet, BMv2 on your Mac or a Linux VM
Free; enough for P2, P3 prototypes
Large-scale simulation
ns-3 or a custom flow-level simulator on a rented multi-core cloud VM
Pay per hour; keep runs scripted
Real GPU traffic traces
A few hours on rented 4β8 GPU cloud instances
Budget carefully; capture traces once and reuse them
Research testbeds
CloudLab, FABRIC, Chameleon (US research testbeds)
Access usually requires an academic affiliation; a collaborator can give you access
Cloud credits
Research credit programs from cloud providers
Usually need a proposal; your research proposal doubles as the application
Budget: plan roughly for cloud spending of a few thousand rupees a month, rising during evaluation phases. Collaborating with an academic lab is the best way to cut this.
Weekly routine and year-end checkpoints
Research needs long, uninterrupted blocks, so this routine puts most hours on weekends and keeps weekdays for reading, writing, and small tasks. Total: about 18β20 hours/week.
Day
Time (IST) Hours
Activity
Monday
9 β 11 PM
2
Read: 1 paper, write its review Tuesday
9 β 11 PM
2
Code: small experiment changes, bug fixes
Wednesday
9 β 10:30 PM
1.5
Write: 500 words on the current paper or thesis Thursday
9 β 11 PM
2
Code or analyze last weekend's results
Friday
β
0
Rest
Saturday
10 AM β 3 PM
5
Deep work: design, large experiments
Sunday
10 AM β 2 PM, 9 β 11 PM
4 + 2
Deep work; then research log + plan next week
Habits that matter more than hours
Keep a dated research log: every experiment, its hypothesis, the result, and what you concluded. This becomes your thesis's raw material.
Start long simulations on Sunday night so they finish during the week.
Monthly: a 30-minute mentor call with a 1-page update sent two days before.
Before each paper deadline, take 1β2 weeks of leave from work if possible; the last weeks before submission are the most intense.
Year-end checkpoints
End of year 1: qualifier passed, survey on arXiv, mentor secured, research proposal approved
End of year 2: P1 submitted, P2 prototype working with first results
End of year 2.5: MS thesis defended; continue-or-stop decision made End of year 3: dissertation proposal defended, P3 submitted
End of year 4: P4 submitted, at least 2 papers accepted overall
End of year 5: P5 done, dissertation defended
Evaluating progress and turning it into a formal degree
Peer review is the real grade: a paper accepted at a top venue is the same evidence whether or not you are enrolled anywhere. Track these signals every 6 months.
Signal
Healthy
Warning
Research log
Entries most weeks
Gaps of a month or more
Submissions
One submission every 6β9 months after year 1
Nothing submitted in 12 months
Reviews received
Reviewers engage with the idea, criticize the evaluation
Reviewers say the problem is not important
Mentor contact
Monthly calls happening
No contact in 3 months
Your motivation
You think about the problem outside scheduled hours
Only deadlines make you work
If reviewers repeatedly say the problem is not important, revisit the agenda with your mentor instead of polishing the same paper. Paths to a formal degree later
Your published work makes these routes much stronger. Rules and eligibility change, so check each institute's current admission pages.
Part-time or external PhD at IITs/IISc for working professionals: some programs let employees register while working, usually with a short coursework requirement and sometimes an employer sponsorship letter. Your papers and a willing faculty co-author make admission far more likely.
PhD by publication: some universities (mainly in the UK and parts of Europe and Australia) award a PhD based on a portfolio of published papers plus a written commentary; check eligibility for external candidates.
Industry research roles: papers at NSDI/SIGCOMM/USENIX Security can qualify you for research-engineer roles at industry labs even without a PhD.
Full-time admission later: if your situation changes, a strong publication record also helps for a funded full-time PhD in India or abroad.
The mentor you find in year 1 matters here too: a faculty member who has co-authored with you is the most natural PhD advisor if you register formally.