# How to be an AI safety research engineer

> Source: <https://www.lesswrong.com/posts/Wty5mcDBEmypap8o2/how-to-be-an-ai-safety-research-engineer>
> Published: 2026-08-10 05:32:00+00:00

This is the advice I wish I had when I started trying to become an AI safety research engineer.

Start by working out which issues you care about. If you don't care about any, hiring managers don't care how good of an engineer you are. You shouldn’t blindly agree with all issues in AI safety. Predicting the future is hard, so many of us will be wrong.

Because everyone is so focused on the shared AI safety mission, people are willing to help you. When entering the field, people will work with you to upskill.

Hence, it's worth being proactive. Email researchers about their papers. But don’t take it personally when someone is too busy to respond. The [AI Safety Map](https://www.aisafety.com/map) gives a good visual overview of who's doing what. Going to conferences and AI safety coworking spaces is particularly important. The field is small, so people who write important papers are often at conferences.

AI safety isn't like medicine. There is no clear path you can slot into and expect to come out with a job. Jobs exist, but you're more likely to find them through people than job boards. Go talk directly to people who have problems they want to solve.

Expect upskilling and job hunting to take about a year. It depends on your background, but it's not going to be quick. Make sure you can afford the career transition.

Remote work is common in AI safety. However, being in a hub means you'll learn a lot more. Actually going to San Francisco or London is really good. I work from Christchurch, but I regularly travel to meet people and learn.

The skills you need depend on the problems you want to solve, so figure that out first. To start, get a good lay of the land on different research agendas. A good place to start is the [Bluedot Technical Alignment course](https://bluedot.org/). You can also read stuff on the [Alignment Forum](https://alignmentforum.org/) to get a range of perspectives.

Python is king. Language models are written in Python. You just have to learn it. Once you've got that down, picking up other languages is helpful.

You also want to get good at deep learning using [PyTorch](https://pytorch.org/). It's hard to do technical work without understanding how deep learning works.

AI safety focuses on language models, so [Andrej Karpathy's "Replicating GPT-2 from scratch"](https://www.youtube.com/watch?v=kCc8FmEb1nY) is an excellent resource for understanding how those work. If you want to go even deeper, the [ARENA curriculum](https://arena.education/) is great. In my eyes, it's the gold standard for being able to do the different bits and pieces within AI safety.

Get good at Git. If you don't know how Git works, you're going to suck as a software engineer. [Docker](https://www.docker.com/) also comes up a lot.

To make or use evals, use [Inspect](https://inspect.ai-safety-institute.org.uk/). For interpretability, [TransformerLens](https://github.com/neelnanda-io/TransformerLens) is essential. So is a deep understanding of transformers.

Linear algebra is essential. Neural nets work by processing a shit ton of numbers. It's really hard to work with language models if you don't know how linear algebra works. [3Blue1Brown’s YouTube channel](https://www.youtube.com/playlist?list=PLZHQObOWTQDPD3MizzM2xVFitgF8hE_ab) is an excellent resource. You will at least want some foundation of probability and calculus. Typically you'd pick that up from university, but online courses work too.

As of writing, skills in research management, operations and evaluations tend to be highly desired

Building good research taste takes lots of time and is genuinely quite difficult. This is not something you learn in a couple of weeks. When you're new, it’s often useful to refer to research ideas from more senior researchers.

Test ideas quickly and move on from ideas that are hard to implement. A good place to start is to find a paper you find interesting, replicate it, and see what problems emerge. For longer projects, start by creating a 1-pager with no references or jargon. Then share it with researchers you trust to refine your ideas. Be aware that it’s common in machine learning that people overstate their results.

It's easy to get stuck down a research direction that just doesn’t matter. Asking yourself "is this overall subfield genuinely a good use of time?" is an important skill. Laying out a theory of change and considering how feasible your work is helps prevent this. Also, avoiding scope creep is needed to ensure you ultimately share your work.

Focus your research on a specific audience who will find your work valuable. Ensure you understand your target audience well. Go speak to the org/person who will use your work. Otherwise, you may create an insightful solution nobody will read. Many papers lie citationless on ArXiv because of this.

Make sure to take ownership of your work. Notice when other areas are not covered by your team and help them. Be willing to adjust as new research questions arise.

People will mainly read your abstract. Ensure you get it right.

It does not matter how good your ideas are unless you can communicate them. Practise writing, public speaking, and the confidence to put ideas out in the world. People get jobs from posting stuff on [LessWrong](https://lesswrong.com/). When people can read your ideas and quickly get intuitions about how you think, it's much easier to work with you.

Default to being careful around controversial takes. Particularly in policy. Sometimes it’s necessary to be controversial, say doing early work on model welfare. Sometimes being controversial will improve your work. But you should ensure your controversial claims are considered before you make them.

Learning alone sucks. Find great people, mentors and communities to learn with. Learning with others keeps you motivated and helps you find future collaborators.

Doing a fellowship is helpful, but it's by no means a silver bullet. Bluedot has around a 30% chance of getting a job after doing the course. Other fellowships tend to be similar. The learnings are still very valuable. Typically, smaller fellowships tend to get higher placement rates. The really famous ones are hard to get into. MATS has a lower acceptance rate than Harvard.

It's important that when you try to get into a fellowship and get rejected, you don't give up. I got rejected from Bluedot initially and now facilitate for them. A rejection does not mean you're unable to do the thing.

Fellowships are much more targeted at getting you good at AI safety than a degree. In my opinion, treat a degree as the second-class option. If you can get into a fellowship, start there. Don’t just consider AI safety fellowships. Start-up incubators or excellent technical programs are often the best step for people.

Use [https://aisafety.com/training](https://aisafety.com/training) for a long list of fellowships. If you need funding to make the transition, Coefficient Giving has [career transition grants](https://op-career-development-funding.paperform.co/). This is often best for mid-career people who have already built valuable career capital.

Doing a degree is much better than nothing. But ideally, try getting into a fellowship first.

One upside of degrees, you can live off your student loan while doing the [ARENA curriculum](https://arena.education/) in the background. Not a bad way to fund your upskilling.

Degrees are reasonably good at upskilling you. But they take ages, and your lecturers won't know everything. You'll be forced to do lectures on stuff that doesn't matter, or spend a lot of time on stuff you don't enjoy.

The reality is degrees are also credentialist. This matters more for more competitive positions. Having really cool papers also serves that function.

Undergrad and master's tend to be more time-efficient than PhDs. If you're doing a PhD and you get the option to do research at a lab doing real things, take that up. Sometimes it’s worth starting, but not finishing a PhD.

In AI safety, it's important to prove that you genuinely care about doing the thing. This helps when someone's trying to employ you because you can turn around and say, "Hey, I've spent lots of hours thinking about this thing, and here are my thoughts to make it better." And then it's so much easier as an employer to justify hiring that person. In my case, that worked through teaching BlueDot's material at my university. BlueDot hired me as I'd already proven that I care about their mission enough to teach the course myself.

Have a bias towards putting your ideas out in the world. People around you can see what you're doing and help you improve. Second, it helps you find employment when people who are hiring read your work.

When you come up with a cool research idea, write it out and put it on [LessWrong](https://lesswrong.com/). When you've got some interesting takes, discuss them with people at a conference.

Go meet people in AI safety. Co-working spaces and conferences are great for this. After a conference, go to co-working spaces in the area.

To get into a co-working space, find the ops person and email them directly. Say something like: "Hey, I deeply care about AI safety, and I'd love to come meet some researchers." People get invited much more than they get declined.

Once you're there, just go talk to researchers over lunch. People are happy to chat, and you can get extremely insightful ideas just from someone's lunch break. When you do go, have something you're working on, a paper you're trying to replicate, a project, whatever. That way you're not floating around wasting people's time, and you can ask targeted questions.

For people in their early career, it’s worth considering a non-ideal role at an ideal org. You can then move internally.

To get into technical fellowships like ARENA or MATS, you often have to do coding interviews. Personally, I kind of hate them because I suck at them. And in practice, people write code using language models all the time now. But you need to be able to write good code *without* one to pass these interviews.

The best place to start is grinding [LeetCode](https://leetcode.com/). It sucks, but it's something you have to do to pass coding interviews. That said, if you don't feel like you're an amazing coder, that doesn't mean you can't be a great researcher. It's just one of those hoops you have to jump through for certain roles.

Not all roles need coding interviews — mine didn't. No need to grind LeetCode until you're applying somewhere that requires it.

If you don't do anything, nothing's going to get done. Go write that fellowship application. Put your ideas down on paper. Learn to program. Replicate a paper. That’s the first step to looking back in a few years and seeing your genuine contribution to making AI safer.
