cd /news/ai-infrastructure/0-mtech-phd-md Β· home β€Ί topics β€Ί ai-infrastructure β€Ί article
[ARTICLE Β· art-142488] src=gist.github.com β†— pub= topic=ai-infrastructure verified=true sentiment=Β· neutral

[0] Mtech_phd.md

A developer known as @danisherror has published a self-directed M.Tech program in Networked Systems & Security for AI Infrastructure, a two-year, four-semester self-study curriculum of roughly 15 hours per week (~1,500 hours total) aimed at building and securing the infrastructure AI runs on, including GPU datacenter networks, sandboxed agent execution, and AI-assisted network automation. The plan sequences six core courses across three semesters plus three electives, running a heavy main track (operating systems, then networks) alongside a lighter side track (Go, then math), with xv6 labs, MIT 6.5840 material, and datacenter papers such as Al-Fares fat-tree, VL2, and Jupiter Rising. It also specifies a weekly schedule, rest weeks every six weeks, and a fallback minimum of about four hours per week when work gets busy.

by read27 min views2 publishedSep 30, 2026

Got it, this is your self-directed M.Tech in Networked Systems & Security for AI Infrastructure. That changes things from my earlier generic answer, so here's a schedule built specifically around this program's Semester 1 (C1 OS, C2 Advanced Networks, C3 Math, C4 Go), fitted to a 5-hour workday.

How much time to use

Your plan is designed for about 15 hours a week. With a 5-hour job you have room for more, but I'd keep 15 hours as the committed load and treat the extra capacity as a buffer (Friday catch-up, longer lab sessions when something breaks). Two years is long, and the plan's own rule of a week off every 6 weeks only works if the weekly load stays sustainable.

The weekly routine (same slots your doc uses; shift them if your work hours differ)

Day

Time

Hrs

What

Mon

9–11 PM

2

Main course: lectures and reading

Tue

9–11 PM

2

Main course: lab coding

Wed

9–10 PM

1

Paper of the week (first and second pass)

Thu

9–11 PM

2

Side course: lectures, exercises

Fri

9–11 PM

0–2

Buffer only: catch up if behind, otherwise rest

Sat

11 AM–2 PM

3

Deep work: main course lab or project

Sun

11 AM–2 PM

3

Side course project or practice

Sun

9–11 PM

2

Paper summary + 15-min weekly review

The key idea is running two tracks at once instead of four courses in parallel: one heavy main course (OS, then Networks) and one lighter side course (Go, then Math). Two-hour sessions split across four courses would mean constant context switching and no lab ever getting finished.

Semester 1 week-by-week (26 weeks, with rest weeks at 7, 14, and 21) Weeks

Main track

Side track

1–6

C1 OS: OSTEP reading + xv6 labs (util + syscall, page tables, traps, COW fork, one lab per week)

Wrap up C3 notes and write the Semester 1 blog post

A few reasons behind the ordering: Go comes first because it's short and you'll need it for the MIT 6.5840 labs in Semester 2. The benchmarking toolkit (C3) is finished by week 16 so it's ready for the C2 congestion-control study. OS comes before Networks because the xv6 networking lab and container work make the TCP stack labs easier to reason about. And your networking background means C2 can move faster than its topic list suggests.

Papers for this semester

Start week 1 with Keshav's "How to Read a Paper," then the three datacenter papers the plan assigns to Semester 1 (Al-Fares fat-tree, VL2, Jupiter Rising) in weeks 2–4. After that, pick papers tied to whatever the main track is covering, such as a classic OS or virtualization paper during C1 and a TCP or datacenter congestion paper during C2.

When work gets busy

Follow the minimum from your plan: Wednesday paper plus one weekend session (about 4 hours), and never skip two weeks in a row. If a course falls more than a couple of weeks behind this table, drop the lowest-value project part and keep the labs, then use the next rest week as a catch-up week instead of a full break.

Want me to add this as a "Semester 1 schedule" tab in your program doc so it sits alongside the course details?

Self-Directed M.Tech: Networked Systems & Security for AI Infrastructure

Sep 30, 2026 Β· @danisherror

Program overview

This is a 2-year, 4-semester self-study program at roughly 15 hours/week (~1,500 hours total), equivalent in depth to an M.Tech in Computer Systems with a research thesis. The goal is to become someone who can build and secure the infrastructure AI runs on: GPU datacenter networks, sandboxed agent execution, and AI-assisted network automation.

Why this track fits you. It builds on what you already do well (C, control planes, IS-IS/SR, TI-LFA/FRR) instead of restarting in crowded pure ML. It connects directly to your MCP approval system, dev-VM sandbox, Process Network Inspector, and the Apstra work at HPE.

How it maps to a real M.Tech.

Real M.Tech component

Equivalent here

Core courses (~6) 6 core courses across Semesters 1–3

Electives (~3) 3 electives in Semester 3, chosen from a list

Lab / assignments

1–2 graded projects per course

Seminar

Paper-reading practice: 1 paper/week, written summaries

Thesis (~1 year) Research project from Semester 3, full-time focus in Semester 4

Degree certificate

Public portfolio: GitHub repos, technical blog, open-source commits, a workshop paper or arXiv preprint Rules that keep it rigorous.

Every course ends with a project that has a written report (2–4 pages) and a public repo.

Use free university courses (MIT, CMU, Stanford, Berkeley) as the lecture backbone; do their labs, not just the videos. Write in C, Go, or Python depending on the layer; C and Go carry most of the systems work.

Course URLs and editions change; search the course code to find the current offering.

Program structure

[embedded content: program roadmap Β· 4 semesters, 3 gates] Semesters 1 and 2 build the systems base; the thesis starts at month 13 beside the AI-layer courses and takes over fully in Semester 4. Do not pass a gate until its condition is met.

Semester 1 (months 1–6): Systems foundations Semester 1 closes the gaps under your networking strength: OS internals, modern networking beyond routing, the math research needs, and Go.

C1. Operating Systems & Systems Programming

Backbone: MIT 6.1810 (formerly 6.S081, xv6 labs); book Operating Systems: Three Easy Pieces (free online). Topics

Processes, threads, scheduling, context switches

Virtual memory, page tables, TLBs, copy-on-write

System calls, traps, interrupts, the kernel/user boundary

Mini-container runtime in C/Go: build a tool that uses namespaces, cgroups v2, pivot_root, seccomp, and dropped capabilities to run a process. Write a threat model listing what it does and does not isolate. This formalizes your recent container investigation.

C2. Advanced Computer Networks

Backbone: Stanford CS144 (build a TCP stack); Computer Networks: A Systems Approach by Peterson & Davie (free online); Stanford CS244 or Princeton COS 561 reading lists for advanced papers.

Topics

TCP internals: congestion control (Reno, CUBIC, BBR), flow control, retransmission

Network measurement and telemetry: sFlow, INT, gNMI streaming

Congestion in datacenters: ECN, DCTCP, incast

Routing at scale: BGP in the datacenter (RFC 7938), EVPN-VXLAN Projects

Complete the CS144 TCP implementation labs.

Congestion-control lab: in Mininet or ns-3, compare CUBIC, BBR, and DCTCP under incast and report the results with graphs.

EVPN-VXLAN fabric in containerlab: build a 2-spine/4-leaf fabric with FRRouting, then break links and measure convergence.

C3. Mathematics for Systems Research

Backbone: MIT 6.041 (Probability), Gilbert Strang's MIT 18.06 (Linear Algebra), selected queueing theory chapters (Mor Harchol-Balter, Performance Modeling and Design of Computer Systems).

Queueing theory: M/M/1, Little's law, load vs latency

Linear algebra essentials for ML (matrices, gradients, tensor shapes)

Graph algorithms: shortest paths, max-flow, spanning trees (you know SPF; add flow and cuts) Statistics for experiments: confidence intervals, variance, how to compare two systems fairly

Project

Statistical benchmarking toolkit (Python): run a workload N times, compute confidence intervals, and draw CDF plots. You will reuse it in every later experiment.

C4. Go for Systems (short course, ~6 weeks)

Backbone:The Go Programming Language (Donovan & Kernighan); Go by Example; Go's official concurrency material.

Topics: goroutines, channels, context cancellation, gRPC/protobuf, profiling with pprof, writing CLIs and daemons.

Project

gNMI telemetry collector in Go: subscribe to streaming telemetry from a containerlab device (SR Linux or cEOS), store it in PostgreSQL or a time-series DB, and alert on link flaps. This maps directly to Apstra device-agent work.

Semester 2 (months 7–12): Core specialization Semester 2 builds the three pillars of the specialization: distributed systems, networks for AI clusters, and systems security.

C5. Distributed Systems

Backbone: MIT 6.5840 (formerly 6.824) with its Go labs; Designing Data-Intensive Applications (Kleppmann). Replicated approval store: rebuild the approval/audit store from your MCP dev-VM project on top of your Raft implementation, so approvals survive node failure and the hash-chained audit log stays consistent.

C6. Datacenter Networking for AI Clusters

This is the course that makes you distinct. No single university course covers it fully, so the backbone is papers plus labs.

Backbone: papers listed under the reading list (Jupiter, fat-tree, DCQCN, RDMA at scale, HPCC, Meta and Alibaba AI-training networks); NVIDIA NCCL documentation; Ultra Ethernet Consortium public material.

Topics

How distributed training works: data/tensor/pipeline parallelism, all-reduce, all-to-all

Collective communication: ring and tree all-reduce, NCCL internals

Failure handling: link flaps, stragglers, how one failure stalls a training job

Fast reroute in the fabric: applying LFA/TI-LFA thinking to AI fabrics

Telemetry and root-cause analysis for training-job slowdowns

Projects

All-reduce simulator (Python or Go): simulate ring vs tree all-reduce on a fat-tree in ns-3 or a custom flow-level simulator; measure job completion time under one failed link. ECMP collision study: show how a few large flows collide under hash-based ECMP and evaluate a flowlet- or spraying-based fix.

Write-up: a 3-page survey "Failure recovery in AI training fabrics" β€” this often becomes the seed of the thesis.

C7. Systems Security

Backbone: Stanford CS155 or Berkeley CS161; MIT 6.5660 (formerly 6.858) labs; Ross Anderson, Security Engineering (3rd ed.). Complete the MIT 6.5660 lab series (privilege separation, web security).

Sandbox comparison study: run the same untrusted workload under Docker, gVisor, Firecracker, and Kata on Linux; compare isolation guarantees, startup time, and syscall overhead with your benchmarking toolkit.

Formal threat model for your MCP CRUD system: attack tree, trust boundaries, and a test suite of malicious-agent scenarios.

Semester 3 (months 13–18): AI systems, programmable networks, electives

Semester 3 adds the AI layer on top of your systems base and starts the thesis in parallel (about 5 of the 15 weekly hours go to research from month 13).

C8. ML Systems (how AI actually runs) Backbone: Andrej Karpathy's Neural Networks: Zero to Hero (build GPT from scratch); CMU 10-414/714 Deep Learning Systems (build a mini PyTorch); Stanford CS336 Language Modeling from Scratch.

Topics

Neural network basics, backpropagation, transformers and attention

Tensors, autograd, GPU memory hierarchy, kernels

Training at scale: data/tensor/pipeline parallelism, ZeRO, checkpointing

Build a small GPT from scratch in PyTorch and train it on a small dataset (your M2 Mac handles small models; use a cloud GPU for a few hours when needed).

Network-trace generator from training: instrument a 2–4 GPU data-parallel job (rented) and capture its all-reduce traffic pattern; feed it into your C6 simulator.

C9. Programmable Networks & Fast Packet Processing

Backbone: p4lang/tutorials on GitHub; Learning eBPF (Liz Rice); DPDK documentation; George Varghese, Network Algorithmics.

Complete the P4 tutorials; implement a P4 program that does per-flow telemetry.

eBPF process-network observer on Linux: a Linux counterpart to your Process Network Inspector that maps sockets and flows to PIDs via eBPF, strictly view-only.

Fast-reroute in P4: implement data-plane link-failure detection and local reroute in BMv2, and compare failover time against control-plane recomputation.

C10. Electives (pick 3) Elective

Backbone

Project

Best if thesis is on

E1. Security of LLM agents

OWASP Top 10 for LLM Applications; papers on prompt injection and CaMeL; AgentDojo benchmark

Red-team your MCP CRUD system with injected tool outputs; measure how many attacks the approval layer blocks

Agent security

E2. Network verification

Batfish docs; papers on header space analysis and config verification

Use Batfish to verify reachability and ECMP properties of your EVPN fabric before and after a change Safe AI network automation

E3. Formal methods (TLA+) Leslie Lamport's TLA+ video course; Specifying Systems

Write a TLA+ spec of your proposal→approval→revalidation→execute protocol and model-check race conditions

Agent security or safe automation

E4. Intent-based networking

Apstra public docs; RFC 9315 (intent-based networking concepts); SONiC docs Build a tiny intent engine: YAML intent β†’ rendered FRR config β†’ Batfish check β†’ deploy to containerlab

Safe AI network automation

E5. Performance engineering

Brendan Gregg's books; Systems Performance

Profile and speed up a real open-source network daemon (e.g., FRR isisd) and upstream a patch

AI fabric reliability

E6. Storage systems

CMU 15-445 (databases); papers on distributed file systems Build a checkpoint store for training jobs and measure recovery time

AI fabric reliability

Recommended default set: E1 + E2 + E3. They cover security, verification, and formal reasoning, which strengthen any of the thesis topics below.

Research projects and thesis options

Pick one thesis topic by month 12 and work on it from month 13 to month 24. The first option is the strongest match for your background; the other three are solid alternatives.

Thesis option A (recommended): Fast failure recovery for AI training fabrics

Research question: When a link or switch fails in a GPU cluster fabric, how much training time is lost, and can local fast-reroute (in the spirit of TI-LFA) plus collective-aware traffic engineering cut that loss significantly?

Why you: almost nobody working on AI networking has shipped TI-LFA in production code. You have.

Method

Build a flow-level simulator of a rail-optimized Clos fabric driven by real all-reduce traces (from C8).

Measure baseline: ECMP rehash plus control-plane reconvergence after failures.

Design a scheme: precomputed backup paths per collective group, placed so the backup avoids congested links.

Implement a prototype in P4 (BMv2) or FRR plus containerlab for the control-plane part.

Evaluate: job completion time, tail latency of collectives, backup-path computation cost, at 64 to 4,096 simulated GPUs.

Target venues: HotNets, the SIGCOMM or NSDI workshops, APNet; arXiv preprint regardless.

Research question: Can an LLM agent propose network changes that are provably safe before deployment, using verification (Batfish), formal intent, and human approval gates?

Method: build a pipeline of natural-language intent β†’ LLM-generated config β†’ Batfish verification β†’ diff-based human approval β†’ staged rollout with automatic rollback. Create a benchmark of 100+ change requests including deliberately risky ones, and measure how many unsafe changes each safety layer catches.

Why you: combines Apstra-style intent networking with your approval-gated MCP design. Directly useful at HPE.

Thesis option C: Verifiable authorization for AI agents acting on real systems

Research question: How can a system prove that every action an agent took was explicitly authorized by a user, even when the agent is manipulated by prompt injection?

Method: formalize your proposal β†’ approval β†’ revalidation β†’ execute protocol in TLA+; implement it with canonical proposal hashes and a transparency-log-style audit; attack it with injected tool outputs (AgentDojo-style); measure attack success rate and approval overhead.

Why you: this is your MCP CRUD and dev-VM work turned into research.

Thesis option D: Sandboxing untrusted agent code

Research question: What is the right isolation boundary (container, gVisor, microVM) for agent-executed code, and what capability model lets agents do useful work without broad access?

Method: systematically compare escape surface, syscall coverage, and startup latency; design an effect-declaration system (from your dev-VM project) and enforce it with seccomp or eBPF LSM.

Smaller research projects (portfolio pieces)

Complete at least two of these alongside the courses; each is 4–8 weeks.

Project

Output

Links to

Measure ECMP imbalance under AI-style elephant flows

Blog post + repo

C6, thesis A

Upstream a fix or feature to FRRouting isisd (e.g., TI-LFA improvements)

Merged patch

Your IS-IS-SR expertise LLM agent that diagnoses a broken containerlab fabric from telemetry

Demo + write-up

C4, C8, thesis B

Prompt-injection test suite for MCP servers

Open-source tool

E1, thesis C

TLA+ spec of an approval protocol, with a found bug documented

Blog post

E3, thesis C

Paper reading list and research skills

Read one paper a week (~100 over two years) and write a 1-page summary for each: problem, key idea, evaluation, weakness, and one follow-up idea. Start with S. Keshav's short guide How to Read a Paper (three-pass method).

Where the field publishes: SIGCOMM, NSDI, HotNets, CoNEXT (networking); OSDI, SOSP, EuroSys, USENIX ATC (systems); USENIX Security, IEEE S&P, CCS (security); MLSys (ML systems). Area

Papers to start with

Read in

Datacenter networks

A Scalable, Commodity Data Center Network Architecture (Al-Fares et al., 2008); VL2 (2009); Jupiter Rising (Google, 2015) Sem 1

Lossless / RDMA networking

Congestion Control for Large-Scale RDMA Deployments (DCQCN, 2015); RDMA over Commodity Ethernet at Scale (2016); HPCC (2019); Swift (2020) Sem 2

AI training networks

RDMA over Ethernet for Distributed AI Training at Meta Scale (2024); Alibaba HPN (2024) Sem 2

Distributed systems

Raft (2014); Paxos Made Simple; MapReduce; GFS; Spanner; Borg Header Space Analysis (2012); A General Approach to Network Configuration Verification (Batfish, 2017)

Sem 3

Isolation

Firecracker (NSDI 2020); gVisor design docs Sem 2

ML systems

Megatron-LM; ZeRO; Efficient Memory Management for LLM Serving with PagedAttention (vLLM, 2023) Sem 3

Agent security

Not what you've signed up for: indirect prompt injection (Greshake et al., 2023); AgentDojo (2024); Defeating Prompt Injections by Design (CaMeL, 2025) Sem 3

Research skills to practice (spread over all four semesters)

Finding a gap: for each paper, write what it did not evaluate.

Writing: follow the structure of a 6-page workshop paper (problem, motivation with numbers, design, evaluation, related work).

Figures: CDFs and time series with clearly labeled axes; reuse your benchmarking toolkit.

Getting feedback: email authors with specific questions, post drafts publicly, ask for reviews in research communities.

Weekly schedule (fits your three-block day) The program needs about 15 hours a week, taken from your 6 PM – 3 AM personal block: 9 hours on weekdays and 6 on weekends. Times are flexible; the weekly hour count is what matters.

Day

Time (IST) Hours

Activity

Monday

9 – 11 PM

2

Lectures and reading for the current course

Tuesday

9 – 11 PM

2

Lab / project coding

Wednesday

9 – 10 PM

1

Paper of the week (first and second pass)

Thursday

9 – 11 PM

2

Lab / project coding

Friday

β€”

0

Rest; buffer for work overruns

Saturday

11 AM – 2 PM

3

Deep work: hardest lab or thesis experiments

Sunday

11 AM – 2 PM, 9 – 11 PM

3 + 2

Deep work; then paper summary + weekly review

Adjustments

In release crunches at work, drop to the minimum: Wednesday paper + one weekend session (~4 hours). Never skip two weeks in a row.

From month 13, Saturday becomes a thesis day. Every 6 weeks, take one full week off. That keeps two years sustainable.

Weekly review (15 minutes, Sunday night)

What did I finish this week?

What is blocked, and why?

Is the current course on track for its 6-month window?

Next week's single most important task

Assessment, milestones, and credibility without a degree

Without a university, your proof is public work. Aim for evidence that a hiring manager or PhD advisor can check in five minutes.

How to grade yourself (per course) Component

Weight

Pass bar

Labs completed

40%

All official labs pass their test suites

Course project

40%

Working repo + 2–4 page report with measured results

Paper summaries

10%

One per week, no gaps longer than 2 weeks

Public write-up 10%

One blog post explaining what you learned

End-of-program portfolio checklist 10 course project repos with READMEs and reports

1 thesis (40–60 pages) and its code, public on GitHub

1 arXiv preprint or workshop paper submission from the thesis

At least 1 merged patch in a real project (FRRouting, containerlab, Batfish, gVisor, or similar)

Self-Directed Research Program: MS by Research and PhD Track

Sep 30, 2026 Β· @danisherror

Program overview

This program trains you to produce publishable research, not to finish courses. It has two tracks: an MS by Research equivalent (about 2.5 years, 1–2 papers and a thesis) and a PhD equivalent (about 5 years total, 3–5 papers and a dissertation). The PhD track continues from the MS track; you decide at the MS gate whether to go on.

How it differs from the M.Tech plan

M.Tech plan

This research program

Main output

Skills + projects

New knowledge: peer-reviewed papers Coursework

~70% of time

~20%, front-loaded in year 1

Success measure

Labs pass, projects work

Reviewers at SIGCOMM/NSDI/USENIX accept the work

Time budget

~15 h/week

~18–20 h/week (research needs long uninterrupted blocks)

Research area:Dependable infrastructure for AI β€” networks and agent-execution systems that fail safely. It combines your control-plane depth (TI-LFA, IS-IS-SR) with your agent-security work (approval-gated MCP, sandboxing).

An honest limit. Self-study cannot grant a degree, and research without feedback drifts. This program therefore builds in external reviewers, collaborators, and real paper submissions from year 1, and ends with a path to register the work for a formal degree (last section).

Both tracks share the first 30 months; at Gate 2 you either stop with a completed MS thesis or continue into the PhD phases.

Phase 1 (months 1–10): Research foundations Phase 1 replaces the coursework year of a research degree: four research-skill modules, a breadth survey, and a qualifying exam you set for yourself and have someone else grade.

R1. Reading and critiquing research (months 1–3)

Method: S. Keshav's three-pass reading; for each paper, write a review in conference format (summary, strengths, weaknesses, questions to authors, accept/reject).

Volume: 3 papers/week for 12 weeks (~36 reviews). Calibration: pick papers from venues that publish reviews or public discussion (e.g., OpenReview) and compare your review with the real ones.

Output: a review notebook on GitHub.

R2. Experimental methodology (months 2–5)

Topics

Choosing baselines that a reviewer will accept

Workloads: real traces vs synthetic; where to find public datacenter and training-job traces

Statistics: confidence intervals, repeated runs, variance sources on shared machines

Simulation vs emulation vs testbed: ns-3, containerlab/Mininet, real hardware; validating a simulator against reality

Reproducibility: artifact packaging, scripts that regenerate every figure

Project: reproduce one published result (for example, a DCQCN or ECMP-imbalance figure) in simulation and write up where your numbers differ and why.

R3. Research writing (months 4–10, continuous)

Study: Simon Peyton Jones's talk How to Write a Great Research Paper; Jim Kurose's advice on writing paper introductions.

Practice: rewrite the introduction of your reproduction report five times, getting feedback each time.

Learn LaTeX and the ACM/USENIX templates; keep a figure-generation pipeline in Python.

R4. Breadth survey (months 5–9) Write a 15–20 page survey: "Failure handling in AI infrastructure: networks, collectives, and agent execution". Cover ~60–80 papers, organize them into a taxonomy, and end with 5–10 open problems. This becomes your map for the next four years and can be posted on arXiv.

Qualifying exam equivalent (month 10) Part

Format

Who grades it

Written

Answer 3 of 5 questions on your survey area, 48 hours, open book

An external mentor (see the advisors section)

Oral

45-minute talk on your survey + open problems, then 30 minutes of questions

Mentor + 1–2 engineers or researchers

Pass bar

You can defend why each open problem matters and how you would evaluate a solution

Mentor's judgment

If you fail, spend 6 more weeks on the weak area and retake it. Do not start the main research agenda before passing. Research agenda

Central question: How do we build AI infrastructure whose failures, whether hardware faults in the network or wrong actions by an AI agent, are contained quickly and provably, without a human debugging every incident?

A PhD is a chain of papers that each answer one part of this question. The first two papers form the MS thesis; papers 3–5 extend it into a PhD.

#

Paper (working title) Research question

Method

Target venue

Track

P1

Measuring the cost of link failures in AI training fabrics

How much job time does one link or switch failure cost at different scales, and where does it go (detection, reconvergence, stragglers)?

Trace-driven simulation of rail-optimized Clos fabrics; validation on a small rented GPU cluster

HotNets or APNet (workshop) MS

P2

Collective-aware fast reroute for GPU fabrics

Can precomputed, collective-aware backup paths (TI-LFA-style) cut the time lost to failures? Can every agent action be tied to an explicit, revalidated user approval, even under prompt injection?

TLA+ protocol spec, implementation from your MCP work, adversarial evaluation

USENIX Security or CCS

PhD

P5

Unified containment for AI infrastructure

Can one framework cover both network faults and agent faults (detect, contain, recover, audit)?

Synthesis of P2–P4 into one system; end-to-end evaluation OSDI or NSDI

PhD

Why this chain works: P1 produces the numbers that motivate P2. P3 and P4 reuse the approval and verification ideas from each other. P5 is the dissertation's unifying contribution. Each paper can stand alone if a later one fails.

Keep a backup topic. If P1 shows that failure cost is already small in practice, pivot to P3/P4 (agent safety) as the main line; the survey from Phase 1 will show other open problems.

MS by Research track (months 11–30) The MS track ends with a thesis built on P1 and P2, at least one submitted paper, and a public defense.

Stages

Months 11–12: Research proposal (5–8 pages). Problem, why it matters (with numbers from your survey), related work, planned approach, evaluation plan, risks. Your mentor must approve it.

Months 13–18: P1 measurement study. Build the simulator, validate it on a small real setup, publish the tool as open source. Submit P1 to a workshop.

Months 19–27: P2 system. Design the fast-reroute scheme, prototype it, evaluate at scale. Submit P2 to a main conference (expect rejection on the first try; revise and resubmit).

Months 28–30: Thesis writing and defense. Write the thesis and give a 45-minute public defense (online meetup, recorded) with your mentor and two outside reviewers asking questions.

Background: AI training traffic, datacenter fabrics, fast reroute (LFA, RLFA, TI-LFA)

Measurement study (P1)

Collective-aware fast reroute: design (P2)

Implementation

Evaluation

Related work

Limitations and future work

Appendix: artifact and reproduction instructions

MS exit criteria

Proposal approved by mentor

P1 submitted (workshop) and posted on arXiv

P2 submitted to a main venue at least once

Code and data public, with a script that regenerates every figure

Thesis written and defended

Decision at month 30: continue to the PhD track only if you enjoyed the uncertainty of research, P2 got substantive reviews (even if rejected), and you still have a question you want to answer. Otherwise, the MS is a strong finish on its own.

PhD track (months 31–60) The PhD track adds P3, P4, and P5 on top of the MS work and ends with a dissertation whose central claim ties them together.

Stages

Months 31–33: Dissertation proposal (15–20 pages). Thesis statement, completed work (P1, P2), planned work (P3–P5), timeline, and risks. Defend it orally before a committee of three: your mentor plus two outside researchers.

Months 34–42: P3, verified AI-generated network change. Build the pipeline and a public benchmark of change requests; submit to NSDI or CoNEXT.

Months 40–50: P4, provable agent authorization. Formal spec in TLA+, implementation, adversarial evaluation; submit to USENIX Security or CCS. Overlaps P3 by design so that a rejection cycle on one keeps the other moving.

Months 50–56: P5, unified containment. The capstone system.

Months 56–60: Dissertation and defense.

Dissertation structure (120–180 pages) Introduction and thesis statement, e.g.: "Failures in AI infrastructure, whether in the network or in agent actions, can be contained within bounded time and with verifiable authorization by combining precomputed local recovery, pre-deployment verification, and explicit approval protocols."

Background and survey (from R4, updated)

Part I: Network failures (P1, P2)

Part II: Agent and automation failures (P3, P4)

Part III: Unified containment (P5)

Related work, limitations, future directions

Publication targets

Level

Count by month 60

Minimum

3 papers accepted, at least 1 at a top venue (SIGCOMM, NSDI, OSDI, SOSP, USENIX Security, IEEE S&P, CCS)

Strong

4–5 accepted, 2+ at top venues, one artifact badge (e.g., artifact evaluation at USENIX)

Also expected

Serve as a reviewer or shadow PC member at least once; 2+ public talks

Rejections are normal: many top-venue papers are rejected once or twice before acceptance. Budget for one resubmission per paper in the timeline.

Advisors, collaborators, compute, and testbeds

The single biggest risk of independent research is working alone. Secure an external mentor by month 6 and a co-author by month 18.

Finding a mentor (the advisor substitute)

From your survey, list 15–20 researchers (faculty at IISc, IITs, or abroad; researchers at Microsoft Research India, Google, Meta, NVIDIA) whose papers you cite most. Email them with something concrete: a reproduction result that differs from their paper, a bug in their artifact, or a short question about their evaluation. Never send a generic "please mentor me" email.

Ask for a light commitment: a 30-minute call once a month.

Expect a low response rate; one yes out of twenty is a good outcome.

Other sources of feedback

Industry colleagues at HPE with networking depth: ask them to review drafts and serve on your "committee".

Research communities and workshops: submit to workshops and student/poster sessions (many systems conferences have them) for early feedback.

Open-source maintainers: FRRouting, containerlab, and Batfish developers can review the practical side of your work.

Apply to research programs that take working professionals as collaborators or visiting researchers; check what Microsoft Research India and IISc/IIT labs currently offer.

Compute and testbeds

Need

Option

Notes

Network emulation

containerlab, Mininet, BMv2 on your Mac or a Linux VM

Free; enough for P2, P3 prototypes

Large-scale simulation

ns-3 or a custom flow-level simulator on a rented multi-core cloud VM

Pay per hour; keep runs scripted

Real GPU traffic traces

A few hours on rented 4–8 GPU cloud instances

Budget carefully; capture traces once and reuse them

Research testbeds

CloudLab, FABRIC, Chameleon (US research testbeds)

Access usually requires an academic affiliation; a collaborator can give you access

Cloud credits

Research credit programs from cloud providers

Usually need a proposal; your research proposal doubles as the application

Budget: plan roughly for cloud spending of a few thousand rupees a month, rising during evaluation phases. Collaborating with an academic lab is the best way to cut this.

Weekly routine and year-end checkpoints

Research needs long, uninterrupted blocks, so this routine puts most hours on weekends and keeps weekdays for reading, writing, and small tasks. Total: about 18–20 hours/week.

Day

Time (IST) Hours

Activity

Monday

9 – 11 PM

2

Read: 1 paper, write its review Tuesday

9 – 11 PM

2

Code: small experiment changes, bug fixes

Wednesday

9 – 10:30 PM

1.5

Write: 500 words on the current paper or thesis Thursday

9 – 11 PM

2

Code or analyze last weekend's results

Friday

β€”

0

Rest

Saturday

10 AM – 3 PM

5

Deep work: design, large experiments

Sunday

10 AM – 2 PM, 9 – 11 PM

4 + 2

Deep work; then research log + plan next week

Habits that matter more than hours

Keep a dated research log: every experiment, its hypothesis, the result, and what you concluded. This becomes your thesis's raw material.

Start long simulations on Sunday night so they finish during the week.

Monthly: a 30-minute mentor call with a 1-page update sent two days before.

Before each paper deadline, take 1–2 weeks of leave from work if possible; the last weeks before submission are the most intense.

Year-end checkpoints

End of year 1: qualifier passed, survey on arXiv, mentor secured, research proposal approved

End of year 2: P1 submitted, P2 prototype working with first results

End of year 2.5: MS thesis defended; continue-or-stop decision made End of year 3: dissertation proposal defended, P3 submitted

End of year 4: P4 submitted, at least 2 papers accepted overall

End of year 5: P5 done, dissertation defended

Evaluating progress and turning it into a formal degree

Peer review is the real grade: a paper accepted at a top venue is the same evidence whether or not you are enrolled anywhere. Track these signals every 6 months.

Signal

Healthy

Warning

Research log

Entries most weeks

Gaps of a month or more

Submissions

One submission every 6–9 months after year 1

Nothing submitted in 12 months

Reviews received

Reviewers engage with the idea, criticize the evaluation

Reviewers say the problem is not important

Mentor contact

Monthly calls happening

No contact in 3 months

Your motivation

You think about the problem outside scheduled hours

Only deadlines make you work

If reviewers repeatedly say the problem is not important, revisit the agenda with your mentor instead of polishing the same paper. Paths to a formal degree later

Your published work makes these routes much stronger. Rules and eligibility change, so check each institute's current admission pages.

Part-time or external PhD at IITs/IISc for working professionals: some programs let employees register while working, usually with a short coursework requirement and sometimes an employer sponsorship letter. Your papers and a willing faculty co-author make admission far more likely.

PhD by publication: some universities (mainly in the UK and parts of Europe and Australia) award a PhD based on a portfolio of published papers plus a written commentary; check eligibility for external candidates.

Industry research roles: papers at NSDI/SIGCOMM/USENIX Security can qualify you for research-engineer roles at industry labs even without a PhD.

Full-time admission later: if your situation changes, a strong publication record also helps for a funded full-time PhD in India or abroad.

The mentor you find in year 1 matters here too: a faculty member who has co-authored with you is the most natural PhD advisor if you register formally.

── more in #ai-infrastructure 4 stories Β· sorted by recency
── more on @@danisherror 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain β€” perfect for shipping the agent you just read about.

$git push zahid main
β†’ Live at https://your-agent.zahid.host βœ“
Get free account β†’ Pricing
from €0/mo Β· no card required
LIVE [news/0-mtech-phd-md] indexed:0 read:27min 2026-09-30 Β· β€”