How many calls I had throughout the process (in addition to normal work).As for companies and interviews, you can see them below (sorted by date). Note that ***ghosted*** means the recruiter never followed up after contact / interviews; ***withdrawn*** means that I politely told the company I was no longer interested (happened only after having a couple offers I was very excited about). The dates collected are **only the interview dates** as I did not have time to scrape my email for a) when I contacted them and b) when I heard back post-interview.
The interview timelines and results for the companies I interviewed with! Salesforce and TRI I exchanged many emails with or even short calls, but I never got into a normal interview pathway. I was not sure what to call them! For the visualizations and data, please see this [repository.](https://github.com/natolambert/job-search-viz)If you want to make these plots even better, you should record a) when you apply / send the email and b) when you hear back. That information was a little too hard to aggregate post-hoc.
I ended up with 4 somewhat different offers (in order of date received). All of the offers I received seemed like very good technical matches and it came down to who I was working with and where my life would be:
- **HuggingFace** -- Research Scientist: an offer to join as the one / one of few RL researchers. I was told I need to take on a role of education my colleagues and collaborating on non-RL projects. This was exciting for me, as I also get to try my hand at what I consider a more stable area of ML research -- NLP.
- **Amazon Robotics & AI** -- Applied Scientist II: an interesting role where I would be doing something along the line of RL / offline RL / continual learning on Amazon's 1000s of deployed fulfillment center robots. I would have learned a lot, but the structures were not clearly in place to continue being a leader in the academic community.
- **Boston Dynamics** -- Research Scientist, Reinforcement Learning: this offer was very exciting. I came close to taking it, but the compensation was weirdly structured. The role would've been leading a new RL team to help translate cutting edge RL research to new applications on the Atlas and Spot robots.
- **NNAISENSE** (pronounced nascence)-- Research Scientist: a company with an impressive amount of expertise in model-based RL, but not a lot of transparency on how the business is doing / growing. At the end of the day I wasn't ready to move full-time to Switzerland, but I really enjoyed all these conversations and would consider joining in the future.
## Job process reflections
In general, the job search process was very empowering for me. If you approach it with an open mind it really can be a high-leverage use of time for the next stage of your career.
- **High yield:** If you're at a good school, the response rate is astounding. I think every company I reached out to (including any job posting I was remotely qualified for online) got back to me with at least an email. I did not need to take the firehouse approach in the slightest. Hopefully, this can let candidates spending more time reflecting and doing interview prep rather than logistical nightmares.****
- **Waiting game:** rejections always will come first. It takes a lot longer to get to the accept phase. There are clearly patterns where you are not the top choice for a company, so you don't hear back for a few weeks. This is okay and expected. Alternatively, sometimes certain companies have total headcount locks mid cycle. It is good to chat with other people in your area who are applying, so you know for sure it is not just you.****
- **Networking value**: AI research is a really small world. The quality of people you get to talk to in this process is amazing. It's a fun process and I think I extracted a ton more long-term value than just my next job.****
- **Research scientist titles**: this title is carrying a lot of prestige for some reason. At larger orgs the differences between engineers and scientists on paper changes promotion paths, compensation, etc., but in terms of the work done they're very similar on moderately sized projects. I think companies like OpenAI have a good title as "Member of Technical Staff" with multiple recruitment pipelines.****
- **Should I do a postdoc?:** there is definitely a good time for a post doc. I don't regret avoiding them, but I do wish I was in a more stable mental place to consider them. I think the default answer to the question of "should I do a postdoc" should be no for people that have industrial inclinations, but... the right postdoc can open so many doors. The right post doc is a safe space to show the world how amazing you are and flesh out your vision. They'll give you no pressure on leaving and many opportunities to grow. I think these are rare, but sound amazing.
### Interview preparation
I bet this is the section where I will get the most interest, but honestly I still don't have any world shattering advice. I can list some of the types of questions I saw a lot, but specific action items for prep are hard to give.
#### Interview styles per **role**
Here's the breakdown of titles I was interviewing for -- 12 Research Scientist, 2 Researcher, 2 ML Engineer, 1 Deep Learning Engineer, 1 RL Engineer, 1 Control Engineer, and 7 unknowns (actually more RS than I thought, some of those really weren't writing papers). Here's some buckets I felt:
- The *paper-writing* **research scientist** interviews: This feels like how the faculty search was described to me. These are all about evaluating your vision, your place in the broader research community, and ability to work with the team.
- The *product-focused* **research scientist** interviews: A mix between faculty search self-promotion and practical engineering questions. Feels like they created this title for the Ph.D. student to feel special, but in reality they will do the same work as other engineers.
- The **machine learning engineer** interviews: For roles without the research angle, I felt like I was pushed more on practical problems. More machine learning fundamentals interviews that I have no idea how I remember any of it (things like logistic regression, probability fundamentals, vector calculus, and math puzzles). Sometimes (e.g. Apple) I had like 3-5 coding interviews. It was so many that I knew I didn't want to do that job.
- The **robotics engineer** interviews: these take me back when I was an electrical engineer and I got fun systems questions. Also *all* of the interviews were extremely specific to robotics and spent less time talking about machine learning or reinforcement learning in general.
#### The most common questions
- What do you want to work on when you show up? Who do you want to work with? If you had unlimited compute what would you work on?
- What was a past project you worked on? What was the impact? What did you learn?
- What styles of communication do you prefer? Do you work in teams? Do you value mentorship?
#### Actionable tips
- If you have the bandwidth, more coding practice will definitely help. I signed up for LeetCode premium for a few months, and this was the best preparation I did. Reading [The Book](https://www.crackingthecodinginterview.com/) is so dry that it's hard to make substantial progress. I did not do that much and made it.
- An increasing number of companies are doing "ML Background" interviews that are mostly math tricks and basic ML tradeoffs. Preparing for this is hard, but studying some coursework would help. I didn't.
- Spend a lot of time on your job talk. It's not about technical content, but conveying a vision and a story. Honestly, the more feedback I got, the fewer plots I felt like I should have in my talk. Magic move is the best transition.
- Write out a research agenda with timelines and goals. It can make you understand what starting a new process will look like.
- You can say no to companies. If they're doing something ridiculous like not communicating the schedule or saying they only have 1 timeslot for an interview, ask them to make space for you. The worst case is they can't accommodate but the vast majority of times they want it to be a positive experience for you.
- Ask what to prepare for interviews. They will often say "leetcode," but sometimes the recruiter actually will give you a topic area like threading or object oriented programing.
### Some benefits of having a job!
It's really amazing having a job. Finishing my ph.d. has given me a little clarity on what waits on the other side for my colleagues who are finishing soon and have been told that it's easier to be happy on the other side.
- **Exit paths:** You can easily change a job, you can't easily move your Ph.D. A Ph.D. is likely the most high stakes job of your life, because if it turns sour you are sacrificing a lot of time to move on. It's wild because most people decide who to work with right after undergrad. My bet would be that people who work for a few years ahead of graduate school have a bit of a better playbook for choosing the right lab. The exit paths also remove the advisor-student power dynamics that plague graduate school -- your manager now needs to invest in you so that they can keep you around and doing your awesome work.
- **A team:** The set of colleagues you have access to is amazing. I joined a research team where I am on the younger side of those with a Ph.D. The breadth of experiences I can easily draw on and engage with makes it a lot easier for me to visualize my next few years and what success looks like. When interviewing this is hard to see, but once you start it should be clear that there are systems in place to help you succeed.
- **Balance:** Work life balance is really strong. Most companies actually respect your life and time.
The HuggingFace Science Team (+ friends)!
## Reinforcement learning reflections
"We solved perception, but our bots keep crashing."
There is a growing optimism -- or maybe more realistically and surprisingly a need for -- reinforcement learning expertise. The most expansive robotic platform in the real-world today is autonomous vehicles. Among these labs, very good perception APIs have been left to control teams. These control engineers have been struggling to develop existing stacks to cover the long-tail of potential problems one encounters when driving in SF. RL practitioners step in with mind- and tool-sets for creating machine-learning based decision making systems.
In many cases I think the technology would involve integrating specific model learning to areas of confusing control performance with high value placed on robustness. With this in mind, most companies happily entertained far out ideas such as "AlphaZero for AVs" or "RL as an adversary for reliability testing." One of my evaluation metrics for a company is how they responded to the question of how they intend on dealing with new and continuous data from partners / deployed robots -- my idea of how all real world data becomes ML.
We'll see in the next few years how these ideas play out, but big players are moving fast in the space. Tesla is growing their RL team, and I do think they generally try and build things that "work" (but often leaves a lot to be desired in terms of evaluating its impact and being transparent on its capabilities). DeepMind seemingly continues to hire everyone I've looked up to that comes onto the job market in RL, so their big projects won't slow at all.
The billion+ dollar question is how big of a data and training scale is needed for something like MuZero to work. We've seen that at maximum scale it is amazing. EfficientZero started to peel that back. So many companies have wanted to try MuZero, so I think we'll know in a couple years.
#### Evaluating RL candidates
I think I've finally reached the chicken and egg problem of a) company wants to hire experts in niche new field, but b) has no one yet so how do you evaluate candidates.
A few of the things I encountered...
- I had a lot of RL interviews, and some of them were suspect. Things like "write down the equation for Q learning" or "implement gradient descent" are normal. Cruise gave me a good reward shaping question (Thanks Alex).
- Tesla also asked a pretty good RL question. It was a big code skeleton to be completed so less pressure on remembering all the details. I think this model where you need to implement a couple lines, understand existing code, and eliminate a bug maps a lot better to real engineering practice.
### Prestige, signal, and noise
A large undercurrent of this post is how helped I am by having the @berkeley.edu email address. My job search would've been very different without this. I think there is an optimum for every candidate on the spectrum of how many emails they should send versus how likely a response is. If you don't expect a response, you probably should spend more time building each on up (by doing more background, building, or networking).
AI is certainly a community driven by prestige. It may not have been a decade ago, but now with the financial upside of it, people are drawn to it for non-scientific reasons. Surely this has impacted me, but I am happy to say that my job decision was not because of the money.
Finally, on noise. This post is **my specific experience.** Even if you are a really similar candidate with similar career goals, this post is already 6months+ out from when I was searching. No process will be repeatable, but the themes will continue to unfold.
For the visualizations and data, please see the accompanying [repository](https://github.com/natolambert/job-search-viz). Thank you to Nicole from [Rora](https://www.teamrora.com/) for your help with this process.
*Thanks to *[*Krishna Murthy*](https://krrish94.github.io/)* for giving me feedback on this post. Thanks to *[*Vitaliy Yevtushenko*](https://www.linkedin.com/in/vital-yevtushenko/)* for suggesting improvements to the plotting code!*
### The last reliable (available) path into AI
URL: https://natolambert.com/writing/path-into-ai
Date: 2022-07-07
Summary: A confluence of trends leaves AI+something, rather than pure AI, as the last great path into machine learning research.
*Notes: v2 3 July 2022 added a lot of text for clarity; a big thank you to Rosanne Liu for providing substantial feedback on this post!*
It seems like every undergrad or graduate school applicant wants to work in artificial intelligence. Though, most people will be throwing their solid applications off to be ignored by only applying to only the best AI labs. I get it, people want to work on AI all the time, but you can still be an AI person if you bring the AI approach to something else. It's an easier path to admission when realize you're passionate about [insert non AI field] and then bring AI into your lab.
## **Context & my terminology**
Throughout this article, I refer to a "pure AI" lab as one predominantly known for their work in AI/ML. While yes Sergey & Pieter work in AI + robotics (something else), the most accessible opportunities are in something further out, where one brings AI in to something new (e.g. physics, energy, analog circuits). Many of my best friends in AI started out in something totally different and made it work. Now, I'm starting to think this path is the most reproducible one in the current AI research landscape.
My hypothesis is: What people should be doing is finding something else that interests them, and applying to do ***AI + something***. This is actually the **TL;DR** of the article -- apply to an applications group using AI as a tool and your odds of acceptance will increase by 10x or more -- so you can read on if you want to learn how UC Berkeley EECS and Berkeley AI Research ([BAIR](https://bair.berkeley.edu/)) got to this point. This post summarizes trends that are happening at all the top CS schools, but I have a much easier time drawing the lines by telling the story I am experiencing.
## **What counts as "pure AI"?**
I know I'm really making this article a little hard to read by tossing around terms like *pure AI* when describing research groups that are doing a lot of applications to AI. I'm trying to distinguish those groups that are known for their work in AI versus those groups that are excellent at another application and may try AI.
Are you considering applying to UC Berkeley? I'll use this as an example. Consider three professors: [Sergey Levine](https://people.eecs.berkeley.edu/~svlevine/), [Claire Tomlin](https://people.eecs.berkeley.edu/~tomlin/), and [Murat Arcak](https://www2.eecs.berkeley.edu/Faculty/Homepages/arcak.html). All of these professors work in fields underpinning recent successes of AI: optimization, robotics, control, etc. How will the number of applicants to each of these groups at Berkeley vary? I would guess that Levine > Tomlin >> Arcak (where each > is an order of magnitude). Though, any savvy student at these groups will have access to the same amazing colleagues and work at Berkeley. If you're worried about getting in, find some "less hyped" professors that match up with problem spaces you're excited about.
For what it's worth, I think most professors listed on the BAIR website are going to get more mentions in their applications. BAIR has an independent admission process, which is definitely more competitive. Before you put all your eggs in that basket, ask people for realistic advice on whether or not you'll get in.
## **The Evolution of Berkeley EECS: The BAIR Takeover**
### 2010
In 2010, the deep learning revolution had not happened (BAIR technically did not exist to my knowledge, but many of the faculty were here doing their thing), and AI researchers and applications of AI had some much smaller minority of the headcount in the department. Consider this drawing below, which is not to scale (there was less AI — like 10-20% in total including applications). Non-AI-identifying was certainly the majority.
**Other AI** is what I refer to professors with a couple students trying out AI applications or AI related fields of control theory, optimization, signal processing, and really anything with strong mathematical foundations. They can call it AI if they want to, which likely was not the case in 2010 but certainly is the case in 2020. Today, what professor in these areas is not going to let a motivated student try and new project with some new ML technique? There are certainly a few, but really graduate students have enough flexibility to go off and do it anyways. The upside is tremendous.
The landscape has really changed in a decade. [BAIR lists](https://bair.berkeley.edu/faculty.html) 63 Faculty, 420 post docs and phd students (lol), and 31 master’s students on their website (Note, some of the members listed *do not* have EECS affiliations, its messy). This is quite the colonization. BAIR lays claim to a large swathe of the faculty and students, many of whom aren't necessarily focused on AI. I count myself as one of the claimed! I’m listed as a BAIR student without an advisor 🤔. Other professors are listed on the BAIR site, yet they regularly spin counter-AI narratives and think some of BAIR’s practices and BAIR’s grandiosity are problematic. Differing opinions — The Berkeley Way.
To put the BAIR numbers in perspective, the [EECS department lists](https://eecs.berkeley.edu/about/by-the-numbers) approximately 190 faculty, 730 graduate students and 3450 undergrads on their website. This puts BAIR at about 30% of the faculty and 50% of the students (astonishing). Adjustment to the first figure is pictured below. The EECS listing has a ton of extra interesting demographics and trends, check it out.
### 2022
How do most people interface with this reality? Through the application process. Don’t apply to pure AI unless you have multiple publications and still are up for a lot of uncertainty. I heard from a reasonably big BAIR group that the collective of BAIR professors will have about applications from 200 top kids and only accept 40. Thousand(s) were likely eliminated before this hat-draw.
This takes us back to my original thesis — apply to AI plus something.
## Admissions
### 2010
If we consider the proportions of faculty I drew above, this could represent the ways faculty got students from the admissions process. Some change of subjects, but all in all it was relatively balanced.
This would not last long.
### 20
As I expect many people in my cohort did, I explored the option of joining a pure AI group after I was admitted and got a respectful *no thank you* from multiple top Berkeley professors. The landscape for who got these opportunities had almost entirely shifted. Though, because the intersections of deep learning with other research topics were, and to a large extent still are, under explored, I got to try fun things in a group not designed around AI research. This journey was largely self-motivated and **opportunistic**.
I went from a normal hardware EE applicant into the pool of people who are trying to figure out how machine learning works. In my first few years, I found a bunch of people doing the same. Year 2-4 of my Ph.D. I knew plenty of people working on robotics, signal processing, optimization, etc. who were not in BAIR but on a more low-key trajectory to the same end goal. The hype in BAIR is a double edged sword for work-life balance and career momentum. Working in applications of AI is generally a less intense field (largely because it is adjacent to the ICML, ICLR, NeurIPs paper-mill) so we got more time to figure out how we work and develop research/scientific fundamentals on top of understanding the numerical shithousery of deep learning.
I see the proportion of admission cases being cognizant of this middle option shrinking while the field of opportunity is still large. Join the alternate AI extravaganza!
### 2022 (Today on)
There is one last fold to the tale that needs to be discussed — the autonomous admission decisions and departmentalizing of BAIR. The figure above does not reflect the true proportions of the admission pool. Those interested in joining the faculty listed on BAIR’s site have only grown in an exponential-like fashion. There is now a clear delineation in application process between those BAIR-targeting students and everyone else. The realistic goal of applications should be to get into Berkeley instead of just BAIR.
In reality, when a student arrives, they can still work with the BAIR professors, they just may need to do a bit more work to do so and find a foundation in another research group. There is an **advantage in a different path** by being able to view problems no one else can.

Stuck on figuring out what to apply AI to and how to fit into Other AI? The point is that these problems are not well defined. Creating a new way to view an application through the lens of deep learning can be a contribution on its own, even if it is not valued by reviewer number two. I encourage people to seek out this perspective and understanding rather than trying to force themselves into a known quantity because of its prestige.
Why is this the last ***reliable* **path? I think it's still hard to execute, but most of the other cards on the table are effectively impossible. I'm excited to hear from those of you who give this a shot -- I believe.
Finally, it's also worth remembering that getting a PhD is not required for most of the industry jobs anymore, but no one I talk to knows a repeatable path to getting them without the credential. You had to be earlier to get in without a Ph.D. Personally, I don't want to rely on luck in getting to where I want to be. I wanted to let everyone know you don't need to too.[ ](https://natolambert.com/writing/path-into-ai#top-im)
### ML/RL & Microrobotics
URL: https://natolambert.com/writing/ml-rl-microrobotics
Date: 2021-09-04
Summary: A memo I wrote to my research group on the open questions when applying machine learning to another research area: novel microrobotics.
I started my Ph.D. thinking I would be working on system's integration and autonomy for microrbotic platforms. A bit of a tangent emerged into model-based reinforcement learning, yet now some more interest is coming to merge these directions within my research group. I wrote a memo on open questions and investigations in this area. I thought some of my followers could be interested in the thought process of applying ML/RL to a novel research area, so I am sharing a lightly edited version here (remove links to open overleaf documents, etc.). Feel free to reach out if you are interested in contributing seriously.
The topics covered in this post follow:
- The motivation of why this is an interesting problem at multiple levels of the stack,
- My contributions, both published and unpublished,
- The many paths we have going forward from here.
## 1. Motivation: We have almost no data, it is hard to understand what we have
Microrobotics, and any Microsystems research, has a very different publication pathway than most of computer science research. The problem is primarily getting something to work once. This can be exemplified in two ways: fabrication / yield problems and testing / experimental problems. The ionocraft has a simple fabrication process but was hard to test (and assemble), but now the MEMs devices for the motors have a more complicated process (due to many, many more moving parts and constraints), but assembly and testing is a bit lower risk.
Therefore, given that one of these data constraints is likely to always exist **we want a method to be extremely sample efficient**. When Googling this, I don’t even think sample efficient is the right word, because we are not doing theoretical work where we *want* it to be more sample efficient, we *need* it to be so, or we will not solve the task. I call this in my work **minimum-data reinforcement learning**.
Second is an unavoidable problem of cutting edge research. With microrobotics, we will likely never have a perfect model of the dynamics. Our simulations will never be perfect. We do not operate like TSMC & ASML. Therefore, we want methods (or students), but hopefully methods, that and learn from data and **reason with uncertainty** to update the design and or control. Progress in machine learning and deep learning particularly has shown a ton of promise in utilizing function approximation for complex tasks and datasets.
Historically, model-based methods have filled the void of sample-efficient and handling uncertainty for people. This is where I started. If you want to watch a talk with a intro to why microrobots, you can see one of my last [practice quals](https://www.youtube.com/watch?v=kA2i0zHWePU) or the beginning of a [seminar](https://www.youtube.com/watch?v=H5Q1UAEkhZY) I gave at Cornell, [slides](https://firebasestorage.googleapis.com/v0/b/firescript-577a2.appspot.com/o/imgs%2Fapp%2Fnatolambert%2FwfaMg3ZYPT.pdf?alt=media&token=8570ad84-8230-44f0-bbf3-c0addd742670).
### More links for microrobotics
I did a quick search for a review paper on recent advancements in microrobotics, but I did not find one. In that stead, I have made a list of some of the robots I know of (feel free to ping me if I should add any!). Many of these are paywalled, let me know if you need access:
- [Ionocraft](https://ieeexplore.ieee.org/stamp/stamp.jsp?arnumber=8373697&casa_token=WKKgH5nHaP0AAAAA:-XBUhgaZ6lGetRQXmFzoC0QB-suTIXfHPSqdE7vAMJspNOXteG1Pkq-gNQr_YAmAtxJX-kq6LcR8&tag=1) (UC Berkeley), [un-named follow up ](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0231362)(U. Washington)
- [Walking microrobots](https://ieeexplore.ieee.org/abstract/document/7994197) (UC Berkeley)
- [Robobee](https://www.jstor.org/stable/pdf/26018027.pdf?casa_token=eVcNrPFlJM0AAAAA:QYltdkiyaKxCM9vSK9HbPDdVRN7PSj-B91eM_QXhpeExdsd9dlnoSc41qbd1L-_o8HqPV6w19TWXdIh2R_rt75E2AWHWVJrD668L8vt0LSWADpVxzPsEtw) (Harvard), [follow-ups](https://www.nature.com/articles/s41586-019-1737-7) are plentiful
- Magnetic-stimulated [swimmer](https://www.liebertpub.com/doi/pdfplus/10.1089/soro.2018.0019) (DGIST & Korean schools)
- [Jumping robot](https://www.proquest.com/openview/195a89ec07b9b65480cbc5afbbca7986/1.pdf/advanced) (UC Berkeley), [laser powered jumper](https://arxiv.org/pdf/1908.03282) (UC Berkeley), [older version](https://ieeexplore.ieee.org/stamp/stamp.jsp?arnumber=4209132&casa_token=0K1TlGYx7ckAAAAA:0VBETFlGKjCQz94xBVuUsxt7NP2ZDNbcTx9rB_BxVBGWBbXfH8R43ZqJWMJPeVfKay17l3Z6uFu0) (UC Berkeley)

Slide from my quals summarizing this problem.
## 2. History: Model-based Reinforcement Learning (MBRL) as a case study
This is the level of the stack where I took a total leap of faith in the spring of 2018. We wanted to control the ionocraft because a) it is super hard to do with any method and b) would be super cool to use a learning method rather than struggling through hand PID tuning. This turned into controlling a quadrotor, which was the closest system we could come up with.
Starting with class projects in [SP18](https://firebasestorage.googleapis.com/v0/b/firescript-577a2.appspot.com/o/imgs%2Fapp%2Fnatolambert%2FTexaD6aywP.pdf?alt=media&token=21bae241-a019-48fe-9927-f5265f9ea78e) (Hybrid Systems with Claire Tomlin and Machine Learning with Anant Sahai) I proceeded to continue working on this over the summer, while helping Dan and Kris with the Ionocraft. In the [FA18](https://firebasestorage.googleapis.com/v0/b/firescript-577a2.appspot.com/o/imgs%2Fapp%2Fnatolambert%2FZ2z4JevNOP.pdf?alt=media&token=9aedef3d-d0aa-4e74-ac75-5f888d254714) Deep RL course, the project was starting to take form. After a [rejected version](https://firebasestorage.googleapis.com/v0/b/firescript-577a2.appspot.com/o/imgs%2Fapp%2Fnatolambert%2Fw3kfx6-Pe4.pdf?alt=media&token=dd47dad1-478e-4805-8659-e2ea931a2d03) at ICRA that year, the final version was [accepted into IROS](https://arxiv.org/pdf/1901.03737)the following spring. So, it took a solid year+ of work to get this out the door (looking at the versions is interesting, but mostly to show the process this can take if you want to take a leap of faith and try to work on this).
Now, there is ***way*** more support of projects in this vein. There is more work on it directly with Kris and the [old guard](https://arxiv.org/abs/2009.01221) and some [M.S. students](https://arxiv.org/abs/2004.13194). There is work I have had generous opportunities to work with Roberto Calandra on external to Berkeley ([example](http://proceedings.mlr.press/v130/zhang21n.html), [example](https://arxiv.org/abs/1909.12324)). Working on this stuff has opened a lot of doors for me, it’s really exciting. There is generally a lack of people a) willing to work on it seriously and b) with a certain niche of skills that makes them useful — for all of you that is easily microrobotics and related EE hardware design.
## 3. Future Work: optimization across the stack of novel robotics
I have ordered these in what I view as most possible in a learning side to least possible. The X-factor of applying these to hardware makes this tricker. For example, doing any control synthesis on hardware is a huge win, and I would **love** to work on this for something like the walking robot. It seems like the timeline for this does not overlap with my critical path, so this may turn into a collaboration at whatever job I get next (hopefully industry research that allows collaborations).
Here is a short summary of problem spaces I discuss:
- **Black box optimization**: data-driven optimization of design where human tuning parameters is hard (high dimensions, nonlinear);
- **Morphology learning**: specific design optimization applied to how a robot will move and or how it operates;
- **Co-adaptation**: jointly optimizing the robot design and the downstream controller for single- or multi-task control;
- **Multi-agent control**: applying our background in swarms to high-dimensional control tasks (kind of a long shot);
- **Controller synthesis**: continuing some of my algorithmic work on model-based RL for controlled generation (kind of a long shot as well);
### Black box design optimization
There is also a white paper I started that someone could take over, and potentially turn into a high-impact paper with **any** real world results ([pdf](https://firebasestorage.googleapis.com/v0/b/firescript-577a2.appspot.com/o/imgs%2Fapp%2Fnatolambert%2FkixcLvDP_f.pdf?alt=media&token=ce1fc017-b014-40f9-babb-85965fb26c17)). It is the idea of using machine learning to optimize design. We have a metric function we want to optimize, this takes the high-dimensional optimization off your hands.
In the past I even made a [GitHub repo](https://github.com/natolambert/mems-bo) to try and make it easier to onboard people here.
Linking to some of the work I have done in model-based RL, there is some [recent](https://arxiv.org/abs/2006.08052) [works ](https://arxiv.org/abs/2107.06882)on model-based optimization of design. Microrobots / real hardware are way more interesting than what has been done, but would likely be worth starting in a simulator. Kris can chime in with some talks he made 20+ years ago on Sugar and accurate MEMs simulation.
**Example deliverable**: Optimize yield of MBRL robots with human-computer joint design optimization.
### Morphology learning & co-adaptation
I broke morphology off into its own category because it is much more focused on locomotion and structural design of how it operates, rather than just optimizing an arbitrary MEMS function (it is a subset of above). Kris has already had multiple generations of undergrads working on things related to this (for example, here is [work ](https://arxiv.org/abs/1803.00196)from Brian Yang and Grant Wang on gaits— they presented it at one of the first group meetings I attended in graduate school). Co-adaptation is when you include the controller design in the problem formulation (a bit more advanced). Roberto has a lot more work here (some with Kris, some ongoing with me and Mark): [example 1](https://arxiv.org/abs/1911.06832) [example 2](https://arxiv.org/abs/1905.01334).
**Example deliverable(s)**: Apply a black-box optimization task to an already existing robot structure (leg length, number of legs, motor force, etc.) and show performance tuning over iterations; Apply an existing MBRL algorithm to joint optimization of design and control on a simulated Microrobot control task.
### Bridging multi-agent learning and hierarchical control
Something that the framing of working on microrobots gives you is a **fundamental expectance of multi-agency**. What I mean here is: anyone working on microrobots bakes into the motivation of their work that eventually we will make a lot of these. Therefore, any method we have to control them must be able to hand high-dimensional input spaces (many agents). There is a fundamental problem in model-based RL methods that they have not solved higher dimension tasks. Therefore, there is potential high impact work with a potentially different approach: break down the complicated control problem for one agent as if different sections are sub-agents in a multi-agent control problem.
This came up when discussion BotNet ([code](https://github.com/PisterLab/BotNet), [paper](https://arxiv.org/abs/2108.13606)) and our continued working group on multi-agent control.
**Example deliverable**: use hierarchical / partially centralized MBRL to control the [Humanoid](https://gym.openai.com/envs/Humanoid-v2/) environment.
### Control synthesis
This work would be addressing many of the open questions in model-based RL research. There is plenty to do, but it may be easier to learn and enter from starting in another area with lower hanging fruit (application work).
**Example deliverable(s)**: apply MBRL to a walking real-world hexapod; developing further sample-efficient MBRL algorithms by considering [objective mismatch](https://arxiv.org/abs/2002.04523).
### Exploitation Exploration (in MBRL)
URL: https://natolambert.com/writing/exploitation-exploration
Date: 2021-06-30
Summary: A few lessons from model-based reinforcement learning how exploration can happen through exploitation of some metric.
Model-based RL does this wonky thing where it explores by letting its controller think it is being successful (planning is actually wrong), but it gets data in a way that works so there hasn’t been huge changes to it. This paradigm of exploring through error, or in MBRL’s case, exploring through model **exploitation**, can come up in any system that doesn’t have an active exploration mechanism built in.
## Review - Classic Exploration Strategies
To start, let’s recap what is generally referred to as the exploration literature. These are all methods where exploration is and **actively designed** part of the system. I quote directly from Lilian Weng who has a great [article](https://lilianweng.github.io/lil-log/2020/06/07/exploration-strategies-in-deep-reinforcement-learning.html) on the *active* exploration mechanisms.
*As a quick recap, let's first go through several classic exploration algorithms that work out pretty well in the multi-armed bandit problem or simple tabular RL.*
*A) ****Epsilon-greedy****: The agent does random exploration occasionally with probability ϵ and takes the optimal action most of the time with probability 1−ϵ.*
*B) ****Upper confidence bounds****: The agent selects the greediest action to maximize the upper confidence bound Q̂ t(a)+Û t(a)Q^t(a)+U^t(a), where Q̂ t(a)Q^t(a) is the average rewards associated with action aa up to time tt and Û t(a)U^t(a) is a function reversely proportional to how many times action aa has been taken. See *[*here*](https://lilianweng.github.io/lil-log/2018/01/23/the-multi-armed-bandit-problem-and-its-solutions.html#upper-confidence-bounds)* for more details.*
*C) ****Boltzmann exploration****: The agent draws actions from a *[*boltzmann distribution*](https://en.wikipedia.org/wiki/Boltzmann_distribution)* (softmax) over the learned Q values, regulated by a temperature parameter ττ.*
*D) ****Thompson sampling****: The agent keeps track of a belief over the probability of optimal actions and samples from this distribution. See *[*here*](https://lilianweng.github.io/lil-log/2018/01/23/the-multi-armed-bandit-problem-and-its-solutions.html#thompson-sampling)* for more details.*
*The following strategies could be used for better exploration in deep RL training when neural networks are used for function approximation:*
*A) ****Entropy loss term****: Add an entropy term H(π(a|s))H(π(a|s)) into the loss function, encouraging the policy to take diverse actions.*
*B) ****Noise-based Exploration****: Add noise into the observation, action or even parameter space (*[*Fortunato, et al. 2017*](https://arxiv.org/abs/1706.10295)*, *[*Plappert, et al. 2017*](https://arxiv.org/abs/1706.01905)*).*
Two newer exploration strategies are detailed in the blog post:
- Count-based exploration: heuristics to track how frequently you visit a state.
- Prediction-based exploration: heuristics to track how your dynamics model (a tool to predict how the environment evolves) performs, in the sense of prediction error, in different regions of the state space. These turn into sort of *error-regularization*prediction mechanisms. They reward systems for new points, but also search for uniformity over model capacity. A version of this happens indirectly in MBRL.
## The Exploration Dance in MBRL
Model-based RL is a dance between getting an expressive model (and the vast amounts of data needed to do so) and having a model that your controller design can work with. Contrary to many optimal control methods, the goal is not to get a *perfect* model, but rather a model that solve your task (or tasks).
Here's what the cartpole task looks like. It's the starting point for many Deep RL experiments -- and easy and interpretable task, keep the pole vertical to live.To a newcomer it may seem like **exploration is not a problem** if the MBRL can solve the task it is trying to do. There is no explicit exploration design involved, so why think about it? I’ve come to learn that MBRL is really **exploitative exploration** because it explores by having the controller think the model is accurate, and taking some random actions averages out over time to a successful controller.
The question with trying to modify the current approach is: *if I try and make the controller know the model is inaccurate (or something similar), does the system lose all of the random exploration it has?*If you change the mechanism for getting the new data, the system may cease to work at all. It’s a chicken-egg problem (all of exploration is, but when you don’t understand the exploration mechanism it is a chicken-egg-knife-edge problem).
A couple things that people have found that are relevant to the exploration question in MBRL:
- model accuracy is proportional to data coverage,
- sampling-based controllers work by averaging out to the correct sequence,
- sampling-based controllers also generate totally new actions by straight up messing up.
### A case study
These lessons are from the appendix of [my paper](https://arxiv.org/pdf/2002.04523.pdf) *Objective Mismatch in Model-based Reinforcement Learning*. Here I show had data distributions can show us the complexity of the exploration problem in MBRL. Essentially we use the PETS optimizer to see how a given dynamics model can solve nearby tasks. By doing this, we see what data was *explored* in the process of solving the task. Below we also look at how the model "accuracy" improves over time (even though it is separate from task performance) and how augmenting a dataset with more random samples can stall the learning process.

Here we use a static dataset (shown below) and study how a controller performs at a moving goal. We want to see how MBRL performance varies w.r.t. the training distribution: in short, performances matches the distribution of data coverage.

The distribution of data collected over cartpole experiments. The data is heavily concentrated around center, and the fall off is rapid, much like control performance.
### Some tools for exploration in MBRL
Here are some thoughts for how exploration could be used as a tool in model-based reinforcement learning:
- **Variance-based control**: If a model has a variance estimate in its forward pass (like many probabilistic ensembles that are used now), then it can incorporate that at control time. It can try to keep a certain amount of uncertainty to the chosen actions are somewhat based on known dynamics. I know a couple people trying to do this, but it hasn’t worked well yet, so maybe the total randomness of sample-based control is the reason MBRL works in that regime.
- **Hybrid training and testing times** (this lends itself nicely to real robots): Have an exploration metric you only use at training time. This is done in some model-free RL, but what is the exploration metric when you are searching for state-action data rather than just looking at an action distribution.
- **Structured exploration** (like entropy-based methods, [SAC](https://arxiv.org/abs/1801.01290)): I think there is a future here in model-based RL, but it may come after some more improvements are made to the model and to the optimizer, as then the system will be a little more *probe-able* (right now MBRL systems are confusing and opaque in parts).
Here is an illustration of how the model accuracy improves over the course of a cartpole learning curve (trials on the bottom, shading is multiple seeds). The model keeps getting more accurate after the task has been solved.
The influence of the number of random points on the learning process. More random data corrupts (slows) the actual learning progress.
##
## Exploration in big, under-designed systems
If big companies deploy learning based systems, especially with RL, the method for exploration will be crucial. The action space may be so broad (amount of content) that no human designer can really understand it. Doubly, truly random actions (as done in epsilon-greedy exploration) may make no sense at all.
I think we may see this exploitation exploration on things like social media when they add RL (because users will complain if the company says *we are showing you truly random content to build our dataset)*. I’m just reluctant because no one has studied it enough to know how it works.
Characterizing exploration in MBRL and other iterative data-building systems is worth studying a bit more.
### Lifelong Learning 2021
URL: https://natolambert.com/writing/lifelong-learning-2021
Date: 2021-06-30
Summary: What I have been learning from recently.
I love to learn. The internet makes it infinitely accessible. These are some things I have learned a lot from.
*Anything on this list I consider very high value for time or cost in the represented area.*
- **P**odcasts: long form podcasts are like being a fly-on-the-wall for fantastic conversations. Learn how to orate and think.
- **B**ooks: books are irreplaceable. If you haven't figured this out yet, keep trying.
- **N**ewsletters: direct-from-source content on various topics. Removes the middleman algorithm that normally delivers your content.
- **T**witter: user-specific communities of thought. Follow the champions of your field(s).
- **A**pps: the internet and smartphones are the ultimate platform where there's always another great app to help you improve.
*Note - I strongly support open podcast RSS feeds. *I am worried about Spotify aggregating podcasts behind a wall, but it's a smart business move.
### Technology
Tech companies are intriguing and hard to predict. I expect to work at tech companies, so it's job training.
- *(N)* [Stratechery ](https://stratechery.com/)from Ben Thompson
- *(P)* [Exponent ](https://exponent.fm/)from Ben Thompson and James Allworth
- *(P)* [Dithering ](https://dithering.fm/)from Ben Thompson and John Gruber
- *(T/N)* [Benedict Evans' Newsletter](https://www.ben-evans.com/)
### Health & Longevity
I am openly obsessed with optimizing my long term health, performance, and mental wellbeing. Investing in one's self is always worth it.
- *(P/N)* [The Drive ](https://peterattiamd.com/podcast/)from Peter Attia
- *(P)* [Endurance Planet](https://www.enduranceplanet.com/) from Tawnee Prazak Gibson
- *(P)* [Found My Fitness](https://www.foundmyfitness.com/) from Rhonda Patrick
- *(A)* [Waking Up](https://wakingup.com/) from Sam Harris
- *(P/B)* [The Ready State](https://thereadystate.com/) from Kelly Starrett
### Artificial Intelligence
I find there is a severe lack of high quality content on AI, automation, and robotics.
- *(P)* [The Artificial Intelligence Podcast](https://lexfridman.com/podcast/) from Lex Friedman
- *(N)* My Newsletter: [Democratizing Automation](https://democraticrobots.substack.com/)? I hope so.
### Life, Current Events, and Philosophy
Filling in the cracks of being a well-rounded human are a bunch of various sources that are teaching me how to think critically, be modest, and treat people equally.
- *(P)* [The Portal ](https://ericweinstein.org/)from Eric Weinstein
- *(P)* [Hardcore History ](https://www.dancarlin.com/)from Dan Carlin
- *(P)* [Making Sensey ](https://samharris.org/podcast/)from Sam Harris
- *(N)* [James Clear's Newsletter](https://jamesclear.com/newsletter)
- *(T/P)* [Naval Ravikant](https://nav.al/subscribe)
If you think there is something I would love, please [send it to me!](mailto:nol@berkeley.edu)
### Robot learning, model-based RL, and related optimization at NeurIPs 2020
URL: https://natolambert.com/writing/neurips-2020
Date: 2021-06-30
Summary: What I learned about deep RL and model-based learning at NeurIPs 2020.
I tried my best to absorb a lot of content at NeurIPs 2020, and it was just as overwhelming as ever. Everyone makes decisions of what content they want to focus on, and it is always an exploration (*learn new things*) versus exploitation (*further mastering material in your area of expertise*) tradeoff. I chose to focus on my areas of expertise: model-based learning, RL, and robotics (I also spent a good bit networking, but that happened in between the lines of these notes).
## Workshops
Some of my mentors have said the workshops are the best parts of conferences, and I am beginning to agree. It is where you see the newest work, honest opinions, and familiar faces (or avatars). The panels this year were where the state of the fields were discussed most openly, so I have compiled my takeaways below.
### Robot Learning Panel
[Link](http://www.robot-learning.ml/2020/) to robot learning workshop. The panelists were Peter Stone (UT Austin), Jeannette Bohg (Stanford University), Dorsa Sadigh (Stanford University), Pete Florence (Google Research, Mountain View), Carolina Parada (Google Research, Mountain View), Jemin Hwangbo (Korea Advanced Institute of Science and Technology), and Fabio Ramos (University of Sydney and NVIDIA).
*** Practical robotics:** The panel began with a discussion as to why we (as academics) have not seen much penetration of learning-based robotic systems in the real world. This is something I have been thinking (and [blogging](https://democraticrobots.substack.com/)) about more frequently — are we actually at an inflection point of robotics in the real world? It’s hard to say. The panelists continued to list some minor applications of learning in real world robotics, but settled on an interesting discussion on the need to **distinguish from industrial applications and industrial consumer products**. I find this interesting — some companies like [Skydio](https://www.skydio.com/) are showing there are some things you can do in the consumer regime. I would like to add that maybe incentives of cemented public companies make the capital cost of adopting robots too hard to warrant on the balance sheets and quarterly reports — when else if not now would robotic cashiers and checkouts become a thing?****
*** End-to-end systems:** There seemed to be a consensus on the robot learning panel that “end to end systems aren't practical”. They spent time discussing the example of navigation (pointing to recent work by Grace Gao) and how classical methods are well better. Carolina from Google was advocating for learning-based systems here in uncertain environments, but using learning for the logistics equivalent of “last mile” delivery seems much better. Make a system that works, then see if continual learning can optimize it over time. The big limits discussed to adoption in industry applications are robustness and safety.
*** Roboticist mentality:** I would like to note from the panel that roboticists seem very fond of their work, network and robots. There was a lengthy discussion on international collaborations and new ways of doing robot experiments during lockdown. This is something I greatly appreciate about the field — people genuinely seem to want to interact with things and show things to work (personally, I am sad I don’t have good systems set up for robotic experiments — this is something I will look for when I go onto the job market).
*** Simulators:** The final point I will make is on a discussion of simulators, broadly encompassing how to define simulators, where to use sim2real, how the definition of a simulator effects the task, and more. The robot learning circles have a very astute sense of models and how they impact results. With no simulator being perfect, it’s interesting to hear about simulator-task matching. Do you spend time trying to make contact forces more accurate or parallelizing the simulator so that you can get more samples?
### Deep RL Panel
[Link](https://sites.google.com/view/deep-rl-workshop-neurips2020/home) to deep reinforcement learning workshop (there was also an [offline RL workshop](https://offline-rl-neurips.github.io/), which seemed interesting. The panelists were Marc Bellemare, Matt Botvinick, Ashley Edwards, Karen Liu, Susan Murphy, Anusha Nagabandi, Pierre-Yves Oudeyer, and Peter Stone.
*** Pace of RL:** The Deep RL panel was very reflective. One of the first questions was trying to tease at if there is a “general slowdown of the pace of the field” and what it means for researchers. Personally, I haven’t seen this. The panelists described it as a slowdown of the pace of breakthroughs, maybe because we have fewer new simulated tasks to solve.****
*** Reproducibility:** Importantly, there was a discussion of hidden elements of paper blocking reproducibility. Essentially, in many RL projects there are many code tricks needed to make it converge, and these don’t end up in papers (at best they are in the appendices). This sucks for reproducibility, but is it more of a competitive race for state-of-the-art? Does it normally matter in practice if our algorithm takes 2x the number of samples to converge? This led into a discussion of incentives.
*** Methods v. Insights:** To quote Anusha, “a good method is fine.” There is an infatuation in the RL community with having insights and good results. Sometimes in research, especially towards applications, a good method is enough of a contribution and making up reasons after the fact for why it is *insightful* may not be the best practice. Anusha commented on how she didn’t have time for insights when she was rushing to try to get the awesome results she did in her PhD. That perspective of just diving into the work and not trying to embellish things can be needed in the field. Ashley Edwards took this one step further to comment that it may be okay if we have fewer people entering the field and fewer eyes following all the work. It removes some of the intensity and may foster creative thinking. I definitely agree — I am trying to come up with problems I think matter and not focusing on citations and paper counts, but it is exhausting.
*** Data v model-structure:** the question that started this discussion was “why has deep RL not had the equivalent of LSTMs, transformers, and CNNs for our field?”. The answer broadly is that RL has an inconsistent data structure and inherently leverages other types of supervised learning. It was interesting to hear them discuss how RL doesn’t have the equivalent of an “imagenet" challenge (MuJoCo does a poor job of doing this because it is hard to work with and expensive), so maybe we aren’t optimizing for general, structural breakthroughs (people make their own problem spaces). An analogy I liked was one of the panelists suggesting maybe researchers should be looking for something like a structured exploration method that could generalize across domains. I am not sure what it looks like, but the ad-hoc-ness of RL is definitely true. The RL framework sort of creates this challenge, but these uncertainties are also why it is so fascinating (harder to define the optimization problem).
*** Where to learn RL:** Interestingly, no one jumped to answer on “where should people go to learn RL?” That clearly is a problem for the field if there are no good links ([I have tried to help](https://natolambert.me/writing/2020/0826_rl.html)!) Eventually they referenced the [RL book](http://incompleteideas.net/book/the-book.html) and discussed how the differing prerequisites to study make it a hard field to start with. For example, many people start from optimal control and Bellman’s principles, but there are also many software-focused CS undergrads trying to dive in and make things work. The differing background and lack of a core “curriculum” makes intra-disciplinary discourse challenging.
## Papers and Leftovers
My [paper on long-term prediction in robots](https://drive.google.com/file/d/1F67sOhbUAy1dc_NigediWXXCQI5HBbPT/view) (or the [video](https://slideslive.com/38941349/learning-accurate-longterm-dynamics-for-modelbased-reinforcement-learning?ref=account-folder-62083-folders)) was well received. Most everyone agrees the current prediction mechanisms are not fantastic, but interestingly most discussions evolve into talking about planning (where sample-based methods are the most common tool). I am glad that I ended up at NeurIPs again, and this was the first year I felt like I was interacting with many people in my field who I had read and heard of, but not really met. Shout out to some colleagues who I enjoyed chatting with: [Michael Zhang (Toronto)](https://michaelrzhang.github.io/), [Oleh Rybkin (Penn)](https://www.seas.upenn.edu/~oleh/), [Thomas Moerland (Delft)](http://thomasmoerland.nl/). I would be happy to try and collaborate with some of these people in the future, there are a lot of overlapping ideas in the community of what should work, but not a lot of certainty on why things do not work yet.
I had flagged some papers before the conference as related to my work, and found some by networking that should be highlighted. A quick note, multiple papers appeared at both the robot learning and deep RL workshops, which I was a little disappointed in. I guess workshops are not monitored and don’t actively prevent that, but it feels slightly disingenuous.
** Multi-Robot Deep Reinforcement Learning via Hierarchically Integrated Models* —[talk](https://slideslive.com/38941325/multirobot-deep-reinforcement-learning-via-hierarchically-integrated-models?ref=account-folder-62083-folders): Hierarchical models (perception and dynamics) combine models with similar video feeds to robots with separate low-level dynamics. This paper was cool because they actually used data from multiple classes of robots.
** Model-based Navigation in Environments with Novel Layouts Using Abstract 2-D Maps* —[paper](https://openreview.net/pdf?id=_lV1OrJIgiG), [talk](https://slideslive.com/38941394/modelbased-navigation-in-environments-with-novel-layouts-using-abstract-2d-maps?ref=account-folder-62083-folders): I found this because we are talking about navigation and multi-agent control a lot more in my group. It’s a little different than what I first thought, but interesting nonetheless.
** Model-Based Reinforcement Learning via Latent-Space Collocation* — [paper](https://drive.google.com/file/d/1zG9NxHxgJFO6Ev1i_bfr-ZeS94uw69oM/view) ,[talk](https://slideslive.com/38941400/modelbased-reinforcement-learning-via-latentspace-collocation?ref=account-folder-62083-folders) : By co-optimizing the state and the actions the agent improves on (visual) model-based rl. The question is: **why hasn’t this worked in state-based RL yet**?
** Continual Model-Based Reinforcement Learning with Hypernetworks* — [paper](https://arxiv.org/abs/2009.11997),[talk](https://slideslive.com/38941346/continual-modelbased-reinforcement-learning-with-hypernetworks?ref=account-folder-62083-folders): Did not have time to read in detail.
** Accelerating Reinforcement Learning with Learned Skill Priors* — [paper](https://arxiv.org/abs/2010.11944),[talk](https://slideslive.com/38941336/accelerating-reinforcement-learning-with-learned-skill-priors?ref=account-folder-62083-folders): This is one of the papers in both Deep RL and Robot Learning workshops.
** Autoregressive Dynamics Models for Offline Policy Evaluation and Optimization*— [paper](https://drive.google.com/file/d/15WjOoXQvx5N8eyFjxJQWCikxT2dI06JV/view): This paper tries to use autoregressive models (predict each state one at a time, allowing the state dimensions to influence each-other and hopefully improve the correlation of model-accuracy to policy improvement). I was super happy to hear from the author it was heavily inspired by some of my past work.
I put talks last because the screenshots take a lot of space.
## Key Talks
There were two keynotes to watch for me.
### [Charles Isbell](https://www.cc.gatech.edu/~isbell/):[You Can’t Escape Hyperparameters and Latent Variables: Machine Learning as a Software Engineering Enterprise](https://nips.cc/virtual/2020/public/invited_16166.html)
This talk is about the scale of the problems we are solving and why they matter. We are compiler hackers, and as a community we need to be SWE, ethnographers, and language nerds. Charles draws many analogies to how we are making hierarchical design decisions, each of which can influence the data flow and bias (much like historical applications of engineering). We need more diverse backgrounds in the loop!
There was an interesting analogy over photography where different products are optimizing for different things that end up being racist. Film was optimized for white colors, and this continues down the line. Software engineering is the practice of translating code to build software **with principles**. Software engineering **is not devoid of bias** and other potential shortcomings. Honestly, just go watch this talk.
### [Marc Deisenroth](https://deisenroth.cc/) and [Cheng Soon Ong](http://www.ong-home.my/):[There and back again, a journey through calculus and gradients](https://slideslive.com/38935796/there-and-back-again-a-tale-of-slopes-and-expectations).
This is a long form tutorial on calculus and linear algebra from the perspective of machine learning. Honestly I loved it just because the authors put in so much effort creating a meme and story from the perspective of middle earth.
## Other Talks
Otherwise, there were some other talks that were interesting, but not “must watch” territory.
### [Besmira Nushi](https://besmiranushi.com/#contact) from MSR on human-centered AI
This talk was focusing on engineering tools to help minimize data and algorithmic bias. I liked parts of it because they actually detailed different design decisions that could be made in realistic scenarios, rather than broadly discussing the problems of AI Bias. They contacted ML engineers rather than just researchers, as they implement everything discussed. The relevant[paper](https://besmiranushi.com/docs/Guidelines-for-Human-AI-Interaction.pdf) and [code](https://github.com/interpretml)
### [Martha White](http://webdocs.cs.ualberta.ca/~whitem/)’s on sources of uncertainty in policy-gradient methods
An interesting [talk](https://slideslive.com/38935821/reducing-variance-and-the-connection-to-valuebased-methods) discussing the three sources of variance in policy-gradient methods of RL (state sampling, action sampling, and reward sampling). It is a very good review of policy gradients, and a lesson for how to investigate a problem space. For state-sampling, people use mini-batches when computing gradients. The other two sources, action and return, involve more detailed solutions. Reducing for actions is by looking at all possible actions (simplest). In practice this doesn't work because it can be expensive and we don't know Q^pi. The solution is to use a baseline. The key is that the control variate z is uncorrelated across the whole expectation of actions.
Why did I include this talk: it is important to be able to reason about your problem space and go into abundant detail about where your implementation may be imperfect. Variance comes up in every numerical/data-driven method because we do not have infinite data. I was making connections during this talk to the problems of model-based learning and how it is very hard to disambiguate uncertainty introduced by the model. In a way, MBRL could be diagnosed like this talk with a fourth source of uncertainty: structural — ie variance from model and controller formulation decisions (this talk focuses on policy gradient specifically).
### America’s Cup boat design via RL
This expo talk from QuantumBlack caught my attention because it combined the[America's cup + RL](https://slideslive.com/38942310/making-boats-fly-by-scaling-reinforcement-learning-with-software-20). I was impressed by the level of implementation details they discussed. There were a lot of items such as starting and stopping a simulation (where PPO is easier to work with than SAC). Also how to deploy many elements and integrate high-fidelity simulations with RL. The crucial question is if this RL actually helps, or if the design is so under-optimized that any offline method would likely give substantial gains.
[I had some more thoughts on twitter](https://twitter.com/natolambert/status/1335620488605937664?s=20).
### Jeff Shamma: Perspectives from Feedback Control on ML [talk](https://neurips.cc/virtual/2020/protected/invited_16168.html)
This talk I included because I think it is an area a lot of people could benefit from, but it requires a lot of skills (machine learning and control theory expertise) to benefit from. It goes through a running example of attitude control of an aircraft and what takeaways people can take from applying learning to classic controls problems.
- Stabilize and shape behavior -- **higher order learning**
- Gradient play (individualized learning in a competitive problem) cannot converge to zero-sum games, as the dynamics of it becomes an unstable system (zeros on diagonal of A matrix).
- Key Idea: see if we can work in a different space of information (such as the history of information) — do this by adding auxiliary states, a common practice in control theory to stabilize a system (like an integrator). Also, see anticipatory learning — adding anticipation allows convergence of nash equilibriums (which makes sense conceptually).
- Robustness to variation -- **passive / monotone learning**
- External and internal parameters can change within the dynamics (parametric) and dynamics change change (e.g. fluids). The dynamic variations introduce new states, potentially an infinite order system. Introduces robust analysis -- how the family of systems performs in context of the controller.
- Example: contractive games where the inner product of the change in strategy and change in payoff -- that is negative (interesting notion of direction). We want the opposite, a passive learning rule, where the payoff and impartial pairwise comparison correlate. We can do robustness analysis for families of systems by abstracting away from a specific controller.
- Track command signals -- **forecasting and no regret learning**
- Jeff breezed through this section, so I didn’t get as much out of it, but it is trying to reason about the need for time separation (and stability w.r.t. a delay) when having one signal tracking another.
### Debugging Deep Model-based Reinforcement Learning Systems
URL: https://natolambert.com/writing/debugging-mbrl
Date: 2021-06-29
Summary: Things I have learned in 3 years of a young, and generally tricky research field.
I saw an [example](https://andyljones.com/posts/rl-debugging.html) of this debugging lessons for model-free RL and felt fairly obliged to repeat it for model-based RL (MBRL). Ultimately MBRL is so much younger and less pervasive, so if I want it to keep growing I need to invest that time in all of you.
For an illustrative case-point, consider these two SOTA codebases:
- [TD3](https://github.com/sfujim/TD3): Twin Delayed Deep Deterministic policy gradient. Reading the code:😀.
- [PETS](https://github.com/kchua/handful-of-trials): Probabilistic Ensembles with Trajectory Sampling. Reading the code: 🤪.
By the title alone, it sounds like they may be equal in complexity, but for model-free algorithms, the modifications take only a few lines of code. In MBRL it takes building entirely new engineering systems (more-or-less). The PETS code has way more moving parts. This post is set up in the following format:
- Overview of model-based RL,
- [Core things](https://natolambert.com/writing/debugging-mbrl#toc-core-tinkering) to tinker with in these systems,
- [Other considerations](https://natolambert.com/writing/debugging-mbrl#toc-other-considerations) that may come up (e.g. when working with robotics),
- [Practical tips](https://natolambert.com/writing/debugging-mbrl#toc-practical-tips-): quick things to change or run for a big potential improvement, and
- [Conclusions](https://natolambert.com/writing/debugging-mbrl#toc-conclusions).
Into the woods we go! This is a brain log of the things I have learned and been stuck on when implementing model-based RL algorithms and applications. To go to the list of practical tips, click [here](https://natolambert.com/writing/debugging-mbrl#toc-practical-tips-).
*Note, want to learn through code? I recommend checking out our *[*MBRL-Lib*](https://github.com/facebookresearch/mbrl-lib)*.*
## MBRL: Overview
Model-based reinforcement learning (MBRL) is an iterative framework for solving tasks in a partially understood environment. There is an agent that repeatedly tries to solve a problem, accumulating state and action data. With that data, the agent creates a structured learning tool — a dynamics model -- to reason about the world. With the dynamics model, the agent decides how to act by predicting into the future. With those actions, the agent collects more data, improves said model, and hopefully improves future actions.
A core framework in a lot of — but certainly not all — recent progress in MBRL is model predictive control (MPC). MPC, which can be optimal, is a control framework for using structured dynamics understanding to solve an optimization when choosing an action. A crucial component of this is the goal of predicting **far into the future**, and deciding in a receding horizon manner. Though, in MBRL this long-term planning is balanced with the understanding that models diverge exponentially as the prediction horizon grows. Many of the debugging tools I discuss are from the lens of long-term planning for control, and could be tweaked to be better phrased for methods using value functions and model-free control.
This post aims to steer clear of any numerical issues in particular, which are research problems in deep learning of some sort, favoring a discussion of trade-offs and weird portions of the system that tend to cause problems (even if we do not know why)!
Secret agent Clank is going to keep improving.
### My biases
I am certainly biased towards deep MBRL for robotics. In the future I see MBRL, and other types of RL being used in many applications, most of which will be digital. I have little experience with CNNs / computer vision / recurrent models, so my advice is very focused on one-step dynamics models. You can see my publications and more obvious biases in my [CV](https://natolambert.com/cv).
## Core Tinkering
Any RL system is primarily determined by the motivation of the tinkerer. The core difference between model-based methods and their model-free brethren is the *addition*of a model. Trying to summarize all the properties of interest for a dynamics model in an iterative, data-driven system [is overwhelming](https://openreview.net/forum?id=p5uylG94S68), and needed. I see the modeling problem as primarily being a *data problem* and a *modeling problem*. The line between these quickly is blurred.
### Dynamics Model Parametrization
How you structure (and gather) your data is of crucial importance.
- **Model type**: the craze is with deep neural networks these days (admittedly the hype is chilling), but there are plenty of other models to consider. Neural network models are becoming easier to use (especially when considering online control) due to the investment in small graphics units, e.g. [Jetson’s](https://developer.nvidia.com/buy-jetson)).
- *Linear models* (e.g. least squares to a state-space system, something like [this](https://www.sciencedirect.com/science/article/pii/S089396591300075X)) are good for elementary systems. It is not quite the least squares approach, but here is a [paper](https://arxiv.org/pdf/1509.06841.pdf) from Professor Levine’s earlier days using a linear model for dynamics.
- *Gaussian Processes* are good when you don’t plan to plan online, have <8 dimensional state-action space, or like structured uncertainty estimates. The slow planning comes from needing to invert a matrix of the number of training points cubed for prediction. I think that if NNs did not exist, GPs would be by far and ahead the leading candidate (PILCO uses them). GPs in this case are heavily limited by a dataset filtering problem to keep the prediction time fast. I would still consider GPs for offline applications though. Also, a tool like [Bayesian Optimization](https://en.wikipedia.org/wiki/Bayesian_optimization) is useful as a companion to a lot of MBRL work (it in itself is a variant of MBRL).
- *Recurrent models* sound fantastic because the goal is to predict the long term future, but… In my experience, LSTMs are not that well suited because they are also hard to train and tend to require more data than control problems provide (I suspect this opinion may change in the coming years). There are two key examples to where LSTMs have been used in MBRL, one outweighing the other (for now). [Dreamer](https://arxiv.org/pdf/1912.01603.pdf) uses an LSTM as a recurrent model for a latent space, and collaborators with Professor Levine [applied collocation](https://openreview.net/forum?id=ku4sJKvnbwV) to a similar latent space. We tried to baselines LSTMs for state-predictions in [this paper](https://arxiv.org/abs/2012.09156), but I think we need some more expertises to make it work (if you have the skills for this, reach out!)
- Use the **delta-state** parametrization: predicting the change in state rather than the true state tends to be better behaved and more useful for control. I have my doubts that it works as well for very unstable (dynamically, think eigenvalues) systems. It is formulated as:
- **Model capacity** has not been pushed to its limit in most tasks. Most papers use canonical values (2 hidden layers, 256 nodes, 5 models in ensemble) and whenever I tune this I see very little impact. My hunch is that the model capacity tends to be way higher than is needed, so most of model training is fighting overfitting with noisy data.
- Advanced model options: append **history** to model input (if the time-constant of the dynamics is well slower than the control frequency actions may take a couple of steps to engage. Increase the models understanding by appending a few past states and actions), **context variables** can be passed into the model as an input and not predicted as an output (e.g. battery voltage, but they [can be misleading](https://firebasestorage.googleapis.com/v0/b/firescript-577a2.appspot.com/o/imgs%2Fapp%2Fnatolambert%2F1uKR9oAgb2.png?alt=media&token=f21f91ea-1bc7-47af-8259-dd8be42b8efa) -- in this case, the load of changing the motor voltage varied the measured battery charge by so much that it was not useful in understand dynamics)
### Dynamics Model Supervised Learning
When the core supervised learning of the model is broken, as in error is not converging with our savior Adam, see many tutorials on this at the base level, but otherwise:
- **Model initialization** can be important. Due to its core use in control, a model that is a bit off can result in performance being stuck at 0. Re-use someone else’s model initialization if you can (it turns out people use [clipped normal distributions](https://www.tensorflow.org/api_docs/python/tf/random/truncated_normal) generally, which is not in PyTorch yet).
- **Anger point**: Some models use **incremental training**, some do not. In the PETS paper, the Half Cheetah results that it touts is using a slightly different model training than vanilla re-training after each trial. In this case, the model parameters are not reinitialized, but rather the optimizer takes more gradient steps from the model parameters used in the previous trial. This results in a slower model change, but it is not well theoretically motivated.
- Need **state/action preprocessing** (e.g. angles to sine and cosine of angle): when planning in systems with angles, if you expect the angles to wrap around (e.g. a heading that keeps turning left will grow from 0 to 2pi to 4pi), the states here need to be processed before labeled to make equivalent angles matched.
### Predicting Multiple Steps
Compounding error emerges when planning multiple steps into the future by accumulating small errors as inputs to the dynamics model. The compounded prediction pass is formulated as:
- **Common bug**: the user does not re-normalize predicted states before passing back into model. Error in this case will blow up really fast. If you want to visualize the trajectories / use them for control (where the whole trajectory matters) prediction usually happens as input -> normalized input -> normalized output -> output -> normalized input …
- Above I mentioned state wrappers for angles, integrating those into predicting trajectories can learn to problems with state/action wrappers. Normally, this is a dimension mismatch in your model (angle becomes sine and cosine of angle), but it can be a pain to implement.
### From Modeling to Planning
Planning usually takes the general form of a finite-horizon model predictive controller:
- Predictive **horizon too short**: some reward functions you won’t even reach without a long enough horizon — when this happens, all candidate actions are equivalent at 0 reward (reduction to random policy)! For example, in a task like manipulation that is completed when an arm gets to a certain epsilon-bubble, with a short model horizon, none of the action sequences may get there.
- Predictive **horizon too long**: best-case point is MBPO, a fun case is PETS. They have an appendix sweeping over horizon (lightly)
- Planning horizons are also not well understood (DeepMind published an entire [paper](https://arxiv.org/pdf/2011.04021.pdf) on this subject)
- **Theoretical limitation**: There is not a signal connecting the measured reward to the dynamics model or optimizer. This is something we call [*Objective Mismatch*](https://arxiv.org/abs/2002.04523), and designing algorithms to address this could be very impactful.
- **Cool option**: If using probabilistic models of some sort, you can plot the model uncertainty at each step in the prediction horizon. Traditionally (not sure why theoretically), the variance diverges to large values or collapses to 0 when the model is outside its training set. This can make for interesting dynamic horizon tuning, but is very hard to design.
### Control Optimization
**Wall clock time** is a huge problem. With all of the planning into the future, MBRL algorithms take a lonnnnnnng time to run. For example, re-producing the PETS experiments on Half Cheetah takes 3 days to run on a GPU. An engineer I worked with literally started experiments then went on holiday time to wait for them. There will be incremental progress improving this, but also maybe someone can build on this [Jax implementation](https://github.com/ikostrikov/jax-rl) that is supposed to offer 2-4x time improvements for some RL algorithms?
- **Experiment carefully**: MBRL is not a research area to just throw tons of experiments at due to the above time limits. Keep track of what you are planning and what you are currently running. Use [Hydra](https://hydra.cc/) to manage your experiments.
- **Weird consensus**: the impression I get from most MBRL researchers using sample-based planning is that the controller works by choosing slightly incorrect actions that over time averages out to a good, but not perfect plan. Understand that current algorithms likely will output weird action sequences if you zoom in on any random seed.
- **Art**: sample-based MPC working on average means you want a model that is very accurate with the best trajectories you have (expert), but also in the nearby area for *robustness*.
### Tuning Reward Functions
MBRL seems to be a bit more closely tied to optimal control methods, and some of the papers formulate the rewards a little more differently (and when you aren’t using the standard baselines, there is more room to tune them)!
- **Practical trend**: data distribution (coverage) is proportional to performance. If you have a environment your MBRL task can solve, you can improve related task performance by getting more labelled data there (think moving the goal state from 0degrees to 15degrees), but you can also limit peak performance by including more random transitions in your dataset.
- If you suspect your reward function is weird, consider **reparametrizing first, tune second**. For example, a cartpole task traditionally gets a living reward for being in a wide range of states (reward is 1 whenever trial is not done), but you can change the behavior to be optimal in the control theory sense by adding a quadratic cost away from the origin.
- **Scalar multiple of rewards suck to tune**: tuning the weights of attitude and trajectory weights is not easy. I hope hierarchical RL takes some of the weight of this.
- **Smooth and bounded rewards** over quadratic costs. For example, taking the cosine of an angle is a nice bounded function that gives a higher value around 0. It drops off much more nicely than the negative of the square.
### Exploration
Gosh, in MBRL exploration happens at the interface between the model and the planner, which makes designing for it very tricky. I wrote a short post on why this is so weird [here](https://natolambert.com/writing/exploitation-exploration). Ultimately, having working MBRL for 3 years now almost full time, exploration is the least-clear path forwards.
To date, exploration seems to happen by chance luck, and integrating more explicit exploration mechanisms into planning (prioritizing an action distribution) could come at the cost of having the wrong model training distribution. I would love to be proven wrong or hear that it is less complicated than I think, of course. A couple of research directions that I feel obliged to share are making improvements on exploration. The first is [plan2explore](https://bair.berkeley.edu/blog/2020/10/06/plan2explore/) (by the author of Dreamer, [*Danijar Hafner*](https://danijar.com/), reminded to me by [Robin Chauhan](https://twitter.com/robinc)):
By maximizing latent disagreement, Plan2Explore selects actions that lead to the largest information gain, therefore improving the model as quickly as possible.
Information gain is a very reasonable approach for the training phase, and it can be turned off at test time (like entropy-based methods in MFRL). A second paper is [*Optimistic Policy Search*](https://arxiv.org/pdf/2006.08684.pdf) in MBRL. These two are likely only scratching the surface, so please let me know if I missed anything.
The missing red links between measured reward and model learning are limiting peak performance of MBRL.
## Other Considerations
There are plenty of things outside the purview of *Deep RL* that affect your system. This is what you learn from trying to advance reinforcement learning from the angle of embodied agents. These practical considerations can have anywhere from 0% to complete-app-breaking influence on your project. When people start trying to productize Deep RL, these will be more main stream.
### System Properties
Some systems are really not meant to be modeled. Deep RL does not discuss things like system eigenvalues, noise, and update rate enough. Ultimately, take a lesson from popular [linear estimators](https://en.wikipedia.org/wiki/Kalman_filter) and know that as you go further into the future with no measurements, the lower bound on model error grows hilariously fast.
### Sensitivity to Hyperparameters
A recent paper I was lucky enough to contribute to changed my mind of the potential of and sensitivity of MBRL algorithms to the **parameters** using them. Ultimately, the core idea is that the best parameters for MBRL may change based on the trial number (e.g. as the agent gets more data, the model gets more accurate, so the predictive horizon can be longer). This is on top of the standard deep learning and reinforcement learning parameter sensitivity.
In general, I would say if you are in a well-supported lab, use an automatic machine learning (AutoML) library, but if you are not, work in an environment that is less competitive (e.g. a real robot you have rather than Mujoco). You can learn more [here](https://arxiv.org/abs/2102.13651).
### Robotics Problems
I am an (potentially inadvisable) unique case where the first time I deployed a MBRL system, it was in the real world. Do not do this, but there are certainly things you can learn about MBRL by using it in the real world that apply in simulation.
- If searching for a stabilizing policy in control, make sure you compare to a **random policy**. When working on [this paper](https://arxiv.org/abs/1901.03737) doing attitude control for a quadrotor, we learned a lot about the fluid dynamics of flight. Long story short: when a quadrotor is close to the ground, the updraft from its propulsion bouncing back creates a little pillow where the robot will effectively have more stable poles (the air both keeps it from sitting on the ground, and makes the pitch-roll unstable pole at 0,0 almost passively stable).
- If your reward function is simple, **correcting exploitative behavior** is not really possible. For example, in the same paper I had the problem (or feature) where the quadrotor realized high thrust was a more stable mode. In this flight mode, it was limited by the height of the room, and progress was saved by me catching the micro-quadrotor as it crashed back towards earth. High-risk, high-fun research got me started here. A more in-depth question is, what part of the model or planner (random sampling in this case!) made it so that behavior only occurred part of the time.
- **Preclusion of application** via computation problems. Because of the system-nature of MBRL with MPC, actually putting this on robots is really hard. It comes down to decisions like: do you run MPC onboard or send the state via radio / wire to another computer to compute the actions. Running controllers off-board makes synchronization, delay, and dropout (communication, not NNs) even more important.
There are no good deep RL libraries designed around real-world systems to data. In the quadrotor project I flew 10 flights of data, transferred it with an usb drive, transferred it back, and ran more experiments. If you look at the learning curve there, getting it was a full day of data collection, training, and hope. You can therefore infer that it obviously took a long time to get that paper done :). A hope I have is that modern open-source projects start giving thought to how their algorithms could be used on real-robots at scale, or maybe that is more of a startup question (product).
- Change your **control frequency**: If you sample states less frequently, the noise in your measurements becomes less of an influence on your labelled training points, which can help your model accuracy. A lower frequency also translates to planning further into the future in time given a set number of steps. Though, for many applications, pushing the control frequency higher lets the behavior correct for bad actions faster.
I call this the "waterfall plot" of candidate actions. Color-coding with respect to predicted reward could be interesting to see if the reward space is smooth with respect to the given action dimension. This was for a quadrotor real-world attitude control task.
## Practical Tips
Here is a list of practical tips for various parts of the puzzle.
### Dynamics Modelling
- Append **history** to model input (if the time-constant of the dynamics is well slower than the control frequency actions may take a couple of steps to engage. Increase the models understanding by appending a few past states and actions).
- **Context variables** can be passed into the model as an input and not predicted as an output (e.g. battery voltage, but they [can be misleading](https://firebasestorage.googleapis.com/v0/b/firescript-577a2.appspot.com/o/imgs%2Fapp%2Fnatolambert%2F1uKR9oAgb2.png?alt=media&token=f21f91ea-1bc7-47af-8259-dd8be42b8efa) -- in this case, the load of changing the motor voltage varied the measured battery charge by so much that it was not useful in understand dynamics).
- Look at the **sorted one-step predictions** in your state-space (Example [here](https://firebasestorage.googleapis.com/v0/b/firescript-577a2.appspot.com/o/imgs%2Fapp%2Fnatolambert%2FUWTBJE_o4y.png?alt=media&token=1f68ef6e-1786-4f2b-983b-216ee31e4c2f)). The points at the edge of the system will tend to look like error is higher, but they should not blow up to infinity.
- Visualize many seeds of predictions and median **prediction error** across multiple horizons. It is likely the case that the bad outliers of prediction may disproportionally harm your performance.
### Control and Planning
- Visualize all of the **candidate trajectories** in a couple dimensions — they should look very diverse.
- Visualize planned trajectories vs ground truth with different replanning frequencies. This will be a series of slowly diverging plans, and it can regularly show the most challenging part of a trajectory by where the plans diverge.
Looking at how closely the predicted trajectories (with chosen actions) match the eventual trajectory is beautiful and useful for considering if your optimizer seems to do anything useful. A lower re-plan frequency here is really just to lower how cluttered your visualization is.
- Check the performance of your agent on the **real dynamics** ([Albert Thomas](https://twitter.com/albertcthomas) reminded me of this one). This takes the dynamics modeling question out of the loop temporarily, but can be very hard to implement. This can be tricky to implement because you need to set the state of the simulator to run your optimizer over it many times and most simulators are not parallelized for GPUs (the best implementation I have used spawned CPU children in Python to accelerate). Note that you do not need to sample over as many actions when doing this because the predictions are actually accurate.
- Don’t actively design an **exploration** mechanism, but visualize how it is working by looking at your dataset over time.
- Something that is not written enough: **cost = -reward**. You can change any cost function to a reward.
## Visualizing MBRL Systems
It is important to visualize what MBRL is doing along some core axes: planning, data accumulation, simulation, and more. Here are some visualizations I have made that helped me learn what is going on.
Dynamics models allow you to simulate many perturbations on the future. As we get better models, this sort of dreaming will be more useful. The Ionocraft is a novel robot in my lab we have been trying to learn to fly with. Let it be said that assembly is a bit too hard for this to be a solo project.
As discussed above, different candidate trajectories will look really different with different reward functions. In this case a "living reward" is a reward when the agent is in an epsilon-bubble of the goal. This is like the default Cartpole reward, but can be extended.
Model based RL ends up getting data close to an expert trajectory -- not data covering the whole state space. Understanding how this dataset looks in different spaces and how it compares to assumptions from optimal control is very important.
## Code examples
Check back soon for a major open-source MBRL project I have been working on. Otherwise, some lighter-weight simulators I have built. A lot has been built on the [original PETS implementation,](https://github.com/kchua/handful-of-trials) but it is in the O.G. TensorFlow.
- A core repository I helped support with Facebook AI is now live [here](https://github.com/facebookresearch/mbrl-lib). The paper describing its design choices can be found [here](https://arxiv.org/abs/2104.10159).
- A good dynamics model can be found [here](https://github.com/natolambert/continuousprediction/blob/master/dynamics_model.py). Why is it useful? It is set up to handle model formulation changes (e.g. true state vs. delta state prediction), normalization, etc. separate to the model type. Also, dynamically making the neural network from a configuration file is needed for any advanced machine learning project.
- A messier repository of many libraries, plotting codes, environments, related optimizers and model simulations, are [here](https://github.com/natolambert/dynamicslearn) (honestly use at your own risk, but browsing for inspiration could be useful).
## Conclusions
I am very optimistic about the future of model-based methods. For some level of research-agenda protection, I don't keep the running list of all the things that I want to do open online (I don't have the people to solve all of them!), yet I am very open to working with new people and talking through these challenges.
Ultimately, when you zoom into any piece of the MBRL puzzle, it is pretty clear that each sub-mechanism is relatively suboptimal. Pushing each piece individually has years of research left in it.
I hope that these pieces start to fit together better, we may get a sort of resonant response, where MBRL unlocks all the things people hope for in it:
- Generalization via a better understanding of the data,
- Interpretability by more optimal controllers with accurate models, and
- Performance by all the pieces gathering.
A good way to look at everything I have public on MBRL is to start from my [writing index](https://natolambert.com/writing) and [Google Scholar](https://scholar.google.com/citations?hl=en&user=O4jW7BsAAAAJ&view_op=list_works&sortby=pubdate).
### Research Progress
My friend [Scott Fujimoto](https://scholar.google.com/citations?user=1Nk3WZoAAAAJ&hl=en) summarizes the state of model-based RL pretty well: it is very intuitive to humans (we plan things out in our head!), but actually trying to implement it is pretty horrendous. I’m hoping to make progress on this.
Some recent papers that you should be aware of in the trajectory of MBRL:
- [PILCO](https://icml.cc/2011/papers/323_icmlpaper.pdf) (2011): Gradient-based policies through a dynamics model.
- [MPPI](https://homes.cs.washington.edu/~bboots/files/InformationTheoreticMPC.pdf) (2017): An alternate model-predictive control (MPC) architecture.
- [PETS](https://arxiv.org/pdf/1805.12114.pdf) (2018) & [POPLIN](https://arxiv.org/pdf/1906.08649.pdf) (2019): Sample-based MPC with weaving trajectories together.
- [Dreamer](https://arxiv.org/pdf/1912.01603.pdf) (2020) & [DreamerV2](https://arxiv.org/pdf/2010.02193.pdf) (2021): Visual MBRL is getting good!
- [MBPO](https://arxiv.org/pdf/1906.08253.pdf) (2019): Model-free RL (SAC) on a learned model.
- [Objective Mismatch](https://arxiv.org/pdf/2002.04523.pdf) (2020): We are starting to understand why MBRL is weird (in theory and numerical measurements).
### Advice
Work on from someone else's implementation when you can, or branch off from their code. I've been around, and helped with, a couple open-source attempts at MBRL code and the pain points are generally things you would not anticipate. These pain points are compounded when you have to debug both the model and planning side -- do yourself a favor and start with one half that you know works.
The field is young: if you think you're onto something, just try it.
Visualize the plans, and actions from models (something model-free cannot do, take advantage of it)! Because the action plans are so crucial to control, learning to whisper with what looks right and what looks long will help you a lot. Below is an example of two dynamics models with the same average accuracy, but one makes for noisy control (from [my paper](https://arxiv.org/abs/2002.04523)). I also think that the tools starting to be built in [this paper](https://openreview.net/forum?id=ZN3s7fN-bo) are compelling for MBRL.
These are two Carpole simulated plans. The right plan is using a model that had its reward lowered with an adversarial attack! The control performance dropped 50% while keeping the model on average "accurate".
### Acknowledgements
Thank you to Luis Pineda for many useful discussions on debugging MBRL while building something exciting. Many of the lessons learned here were with Roberto Calandra, and some with Daniel Drew and Omry Yadan. Thank you to Eugene Vinitsky for feedback on my first draft.
If you feel the need, I have made a citation for this one.
Thanks for reading! If you have a question, ping me on [Twitter](https://twitter.com/natolambert).
### All grad students (should) study graphic design
URL: https://natolambert.com/writing/grad-students-graphic-design
Date: 2021-06-28
Summary: People judge your papers by their cover. You can trick them into believing your science with pretty pictures.
Tricks I have learned, code I have written, and free help to make your manuscripts better.
I wrote a [limerick](https://en.wikipedia.org/wiki/Limerick_%28poetry%29) describing the experience every graduate student gets to.
*Reviewers don’t even read the text,*
*So making the plots pretty seems like a hex.*
*But the preprints I’ve seen,*
*So seldom are clean,*
*And the clean ones increase your h-index.*
If [Twitter announcements increase the citations of work](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0229446), making the content visually pleasing has to do even more. am not even entirely convinced that the figures are the single important part of making a paper aesthetically pleasing, but making the whole package right is crucial for acceptance and for people to cite your work. The [science of first-impressions](https://www.psychologicalscience.org/observer/studying-first-impressions-what-to-consider) definitely goes beyond just people and into work, websites, and all of the content we consume.
This post is half brain-log of the tools I have made during my PhD and half an ode-to design every graduate student should keep track of.
This is the default [matplotlib](https://matplotlib.org/examples/pylab_examples/simple_plot.html) example. The colors alone are not bad, but when you add more lines they get strange. Honestly, over time the simplicity of it has grown on me, but there is a bias that people have around matplotlib plots looking preliminary. Refining the color, style, and shape of your figures makes them more memorable.
Having a unique plot scheme is like having a unique website. The first time someone sees it they will think “oh this is cool” and from the second time onwards they will say “I trust this work”. It’s a sort of research-brand. When using matplotlib basics you are tying your brand to whatever papers other people have read with the same default plots (there is definitely a distribution of truth to this, but I don’t see it as something worth risking).
## Make plotting primitives (I use plotly, mostly)
Figuring out what the hell is going on with python plots takes a hilarious amount of time. I have ad-hoc built a set of plotting functions I now use in every project (mostly for better and sometimes for worse).
You can find my code [here](https://github.com/natolambert/plotting-basics). Two notable things I have added to plotly (that matplotlib does better by default):
- **Adding markers every *n* points**: when you have a lot of data, you don’t want markers on every point! Doubly, you want your markers to be big and distinct. Big and distinct doesn’t work well when they overlap (*note*: my function also can randomize the start point so the markers are spread out.).
- **Error bar line plot**: some reason plotly doesn’t have error bar line plots (shaded regions). I modified code I found online to do proper error bar plots (set the quartiles you want to visualize). The different ways of doing error bars is interesting. On one hand there is mean with min and max, but that is very prone to outliers. Median with 65th and 95th percentiles tends to do pretty well!

One of my favorite plots of all time. Studying an advarsarial attack in my paper "Objective Mismatch in Model-based Reinforcement Learning."
### Photoshop for your plots: Save as pdf, edit as pdf
When fine tuning a camera ready figure, open it in something like Adobe Illustrator and move the pieces around to make them look right. If you have a lot of elements you can use select > select same color to change a set of pieces at the same time. It can be easier to make your own axis labels in after effects rather than through code. This is effectively photoshop for your plots (other vector formats work too).
### Choosing colors

Conversations with myself regularly are testing colors.You can use slackbot to evaluate color options. Low-contrast colors tend to be easier on the eyes. Feel free to use the four below that I originally got from [Ge Yang](https://twitter.com/episodeyang?lang=en).
Use [alpha](https://en.wikipedia.org/wiki/Alpha_compositing) when a plot has a lot going on. This makes the points slightly see through and adds a sense of depth to a 2 dimensional plane.
## LaTeX wizardy
Not all of making a paper pretty falls into the rectangles of color. Integrating the text around the figures is important, and can take more time. This is the part of paper writing that feels like a massage.
### Figure prep
A series of things should be done to your plot to make it ready for in paper:
- White background,
- Remove top & right axis lines (maybe more),
- No gridlines,
- Match font to manuscript,
- Configure aspect ratio,
- setting the axes is the only data manipulation you get! Use it!
{{ - - - - - CODE - Python - - - - - }}
fig_plo.update_xaxes(title_text="Year", linecolor='black', # account for white background row=1, col=1, zeroline=True, zerolinecolor='rgba(0,0,0,.5)', zerolinewidth=1,) fig_plo.update_yaxes(title_text="Market Share (%)", linecolor='black', # account for white background row=1, col=1, zeroline=True, zerolinecolor='rgba(0,0,0,.5)', zerolinewidth=1,)
{{- - - - - ENDCODE - - - - - }}
### *LaTeX* Legends
Never worry about getting the legend perfect anymore, make it in the manuscript! I don’t remember where I picked this up, but I no longer make my legends in code.
{{- - - - - CODE - Tex - - - - - }}

I use this type of LaTeX legend in most of my papers. It is easy to read and reduces code burden.
### Making your text look right
Some tips for things to avoid in your manuscript:
- hanging lines in section titles,
- wide variance of paragraph lengths,
- figures coming well after their reference in text,
- citations can be wrong — have a consistent format of venue (either with or without the conference abbreviation).
You can download an example preamble for LaTeX [here](https://firebasestorage.googleapis.com/v0/b/firescript-577a2.appspot.com/o/imgs%2Fapp%2Fnatolambert%2FW0BZK-12x1.tex?alt=media&token=f2400ce6-41cf-4458-beb5-f18e554ce658). It has clean ways to reference equations, sections, etc. and basic math operators like absolute value, vectors, and things to make your pdf nicer.
Colors! Nice!
## Wrap up
Writing papers is really a privilege (as long as you are not forcing it out the door) — it means you have done the hard work to create new knowledge, and now you get to craft it into a story. An older version of this post can be [found on Medium](https://towardsdatascience.com/crisp-python-plots-based-on-visualization-theory-5ac3a82c398e).
I love the graphic design side of things. This doesn’t mean you need to. You need to try and make your work pleasing if you want it to be received well.
### Reflecting on being a graduate student (in AI) in 2020
URL: https://natolambert.com/writing/reflecting-on-being-a-graduate-student-in-ai-in-2020
Date: 2021-04-05
Summary: Starting to build my guide and advice for graduate school.
Being a graduate student is uniquely hard, but also in ways that are hard to describe. I am specifically thinking about PhD students here, but it applies to everyone to varying degrees. Why is it that I hear things from recent graduates like “there is light at the end of the tunnel” or the simple “you can do it” whenever discussing my graduation date coming up.
## The broken power dynamics of a PhD program
**There is something uniquely challenging about getting through a PhD**. The fun fact with most of the people who say this is that the work they are doing often doesn’t change much when they leave, it is just a change in compensation, incentives, environment, and mindset.
- **Compensation**: being paid the actual market rate for your work removes a lot of anxiety
- **Incentives**: even though in industry research labs the goal may be to publish papers, there are normally far fewer webs saying things like “your papers need to form a coherent story” or “you need 3 journal papers before you can get promoted” — research is a part of a story and the incentives are better matched to that. The number one incentive of a PhD is *convice my advisor to let me leave*.
- **Environment**: graduate students almost always have shitty desk space, not great home space, not great separation of work and life. These things add up in so many ways. It is like accumulated advantage of going to good schools, having a good brand, etc., but it slowly pulls you down and weighs on people: *accumulated baggage* (colloquially known as jadedness).
- **Mindset**: this is pretty much the intangible part of being a graduate student. I think the reality of many graduate schools is ***better than*** people give it credit for. The collective mindset of graduate students and their expectations can *pull each-other down to their expectations*.
I have spent so much time trying to break narratives (both in my head and my friends) of “I need to be working, my advisor is in control, my advisor knows what is best, I am inadequate, etc.” It is so sickening, but real, and we need to consider how it propagates to help eliminate it. Generally, PhD program renovations (think improvements in overall wellbeing) fall short due to the dramatic power imbalance between advisor and advisee, and how the department has no cards to adjust it. Think of if a student filed an anonymous complaint about their advisor because the number of students is so low and doing things like getting a new job (lab) is near-impossible, the advisor would know who it is and the advisee has no way out. Ultimately, the only way to reform these dynamics is to have both **better informed** and **more caring** professors (ignore the problem of tenure, aging professorships, and research-based metrics for faculty jobs).
**Navigating the advisor-advisee dynamics and placing yourself in the right research group is the key to graduate school**. It’s a shame that lab’s and advisors come in packages because some labs would work great for people, but the advisor does not.
### The context of AI research
With everything I say below, I see the AI field being amplified in almost all the challenges. AI is so lucrative and growing at a near exponential rate, so the levels of competition are naturally rising. Though, **academia already is competitive enough, so piling on top of this is not a recipe for success** for the field of AI itself. The level of excitement around AI is not the norm across research. I started my career in [MEMs](https://en.wikipedia.org/wiki/Microelectromechanical_systems) and slowly made my space in AI. It’s a mess and having one foot out of the AI pool is refreshing.
## Challenges I (and all of us) have dealt with
Here is a list of things I have had to deal with because most people won’t tell you what goes wrong in their PhD, even though everyone has them.
- Failed the preliminary exam on my first try,
- Somewhere from 2-4 paper rejections on the first submission (do not want to count),
- One paper rejected after being recommended for publication by the area chair, and we still don’t know why,
- Teaching two extra semesters because funding fell through.
These are all normal things. I am happy to be making it through the program and accept it will not be perfect. Here’s the first photo I took at Berkeley. I am now going to go through some quicker thoughts.
## Building models from incomplete information, a masterclass in type A stress mongering
There are definitely many types of people who go on to get a PhD, but with how competitive top programs have become it is becoming increasingly packed with type A students who started their research career early and are trying to make a big name for themselves (having more of these students is bad for intellectual diversity in my opinion).
### Papers, publications, and citations
It’s slightly amazing to me that somehow most people starting (and applying for) graduate school don’t know the real process to get a paper accepted and how random it is. Multiple studies show that the peer review process at top AI conferences (see [this one](http://blog.mrtz.org/2014/12/15/the-nips-experiment.html)) overwhelmingly random. The TL;DR of what people should know:
- Papers aren’t publications: a paper is written and published at different times. See below.
- Acceptances are random: a good paper can be rejected multiple times and weird papers can win awards. It can take a year to be officially accepted and this can have a big impact on an applicant even though the research they did does not really change.
- Citations are more random: people often choose what papers to cite by a Google query — this is not a rigorous system. People often cite so much of their work and that can *bootstrap a perceived as famous paper*.
- Citations are field specific: different venues have different rules on how many citations are allowed per paper and what constitutes a paper. Fields like computer vision are known for having tons of citations.
I still don’t really understand these mechanics, yet they influence how I value myself sometimes. That is very silly and we should not do it.
### Comparing metrics
The biggest challenge for me in 2020 has been stopping tracking how many papers I have in the pipeline, if I will have enough citations to get a job I want, and more forms of this. Ultimately, comparing yourself as a graduate student to other students is ***a total waste of time***. Each student is one member of a small research group with vastly different circumstances.
**Regardless, research and citations are not a zero-sum game, we should want everyone to succeed!**
**Many projects fail**, so many students don’t publish all of their work. This random process over a few years of graduate study creates a huge amount of variation.
Some professors literally take years off from having students. The students that join to start the new group have to re-establish how to write a paper. Some other students are added to papers as a first year for watching experiments. The constant here is learning how to do research, not production. *Citations are the observation we get from a very complicated system.* Making any conclusions on the worth of a student from this point of view is sloppy and propagating problematic assumptions.
## People don’t talk about the student part of graduate school
A big part of why PhD programs are harder than jobs people get is because PhD students have to do just that, be a student. Classes, teaching, operating at a bureaucratic university, and all take up anywhere from 10-40 hours a week in a given semester. We are expected to fit in a full-time research career moreover.
Not only do the hours do not add up, but the expectations around it and the mental strain cannot be ignored.
### Teaching, classes, and unclear expectations
Teaching is one of the most intellectually demanding activities I have done (with pure research being near behind). No one can do these for more than 4 hours a day sustainably (the people who can do more than 4 are the super famous ones). Teaching can easily take up multiple days a week, but there are no criteria in the programs other than needing to do them to graduate.
It is often said that teaching experience can help get jobs. I did a survey of people I know in industry and 0 of them said teaching is something people directly look for, but rather **traits that make someone a useful instructor often correlate successfully in many jobs**. Comically, while teaching experience is *useful* when applying to faculty experience, it can be ignored in cases of some high-profile hires.
Finally, we have to navigate teaching and classes within the scope of our advisor. Some advisors ignore that these things exist and expect the same research output, while some encourage it but then it ultimately delays students graduating for putting off their core research. Graduate students do deserve better, but being careful with these dynamics is important.
### Finding funding
Every student has their own funding journey (expect those joining the most famous labs, their journey was starting their hard work and being lucky early). I have written about 10 grants in my PhD and even those students with funding have to do funding reviews for work they do not necessarily care about. The research vector needed for funding and the looming dark cloud of needing to look for funding every semester tick like clockwork for some students. Not knowing how you will be paid your below market rate wage a few times a year can add an unexpected chunk of stress.
### 2020 and graduate school
This year was a big test. I think we are making it through, but it is obvious that most departments are not set up to handle something like this. Admissions were delayed, funding fell out, new students started remote, along with a litany of other new problems. I’m proud of my colleagues for making it through the year, but I am acknowledging that many people had to leave their programs. In a way graduate school is a privilege, but it should also do more to be accessible to all.
## Backcasting: why you should not totally believe my advice
Ultimately, every time you get advice from someone it is what worked for them. What works to make one person super famous may cause harm on average across a field. Searching for and acting on advice is a slow, iterative, and reflective process.
I am happy to add more advice to these lists. Some of the more benign ones seem obvious, but they aren’t commonly practiced, so they need to be said more.
### Advice I have stuck with
The best advice I received when started was “a PhD is a marathon not a sprint.” Ultimately, it is okay to be down, it is okay to have slow days, and it is okay to mess up, as long as you keep trying. Everyone burns out during their PhD, and being primed by this innocuous quote has made my recovery a little easier. Some other advice is:
- Contacting graduate students is easier than professors.
- Be positive and engage with people and opportunities will come of it.
- Have work-life balance and take care of your body.
### Advice that may be good
This is mostly things that I wish I did and I am trying to figure out how to phrase:
- Focus on **thinking clearly and making routines that work for you**, rather than searching to please other people.
- Don’t undervalue the importance of finding the right advisor.
- Leaving with a masters is not shameful: there are so few lab positions available, that the one right for you may not have had funding. That’s unlucky, but okay.
- Take more time off now, don't wait until you've "made it" or anything. 1. builds better habits for taking time off on holidays as you get older and...2. when you have more paper / review / misc deadlines it is so much harder.
### Advice that will only work out if you are lucky
All the advice below I don’t see working because ultimately a PhD is too short to really know how the numbers work out.
- Maximize the number of papers you are on.
- Optimize for citations.
- Choose an advisor because they are famous.
## Closing thoughts
A PhD should teach you how to create knowledge in a new area, and everything that happens along the way is secondary to its goals (yes, that includes papers). It is better to graduate a good researcher than as a researcher with a couple of good papers you got your name attached to. If you continue in academia, the time in your PhD will be one small segment of your career and hopefully the start of some compounding growth. If you leave academia, the skills you acquire along the way are generally all that matters anyways.
I am very lucky to be working in a field that is relevant to our short-term future (AI). I did this somewhat intentionally by listening for opportunities to try it. If you are interested in AI and want to do research in it, the best bet may be to have other interests and expertises as well.
If you are starting your journey in graduate school, good luck, you **can do it**, and your path is valid and impressive.
### A Different Intro to RL in 30 Minutes
URL: https://natolambert.com/writing/intro-to-rl-in-30-minutes
Date: 2021-02-16
Summary: A 30 minute conceptual intro to Markov decision processes, iterative updates, and reinforcement learning.
This document was originally a collection of blog posts with thousands of views on Medium, but honestly Medium doesn’t deserve them (clickbait eats quality there, and the whims of their editors). You’ll get an introduction to reinforcement learning how I learned to teach them at UC Berkeley for CS 188: Intro to Artificial Intelligence. I take a big detour into linear algebra to help ground some of the material in a field with a wider application base.
## Brief History
Reinforcement learning holds its roots in the history of optimal control. The story began in the 1950s with exact dynamic programming, which broadly speaking is the structured approach of breaking down a confined problem into smaller, solvable sub-problems [wikipedia](https://en.wikipedia.org/wiki/Dynamic_programming), credited to [Richard Bellman](https://natolambert.me/writing/2020/(https://en.wikipedia.org/wiki/Richard_E._Bellman). Good history to know is that Claude Shannon and Richard Bellman revolutionized many computational sciences in the 1950s and 1960s.
Through the 1980s, some initial work on the link between RL and control emerged, and the first notable result was the [Backgammon programs of Tesauro](https://en.wikipedia.org/wiki/TD-Gammon) based on temporal-difference models in 1992. Through the 1990s, more analysis of algorithms emerged and leaned towards what we now call RL. A seminal paper is “Simple Statistical Gradient-Following Algorithms for Connectionist Reinforcement Learning” from Ronald J. Williams, which introduced what is now vanilla policy gradient. Note that in the title he included the term ‘Connectionist’ to describe RL — this was his way of specifying his algorithm towards models following the design of human cognition. These are now called neural networks, but just two and a half decades ago was a small subfield of investigation.
It was not until the mid-2000s, with the advent of big data and the computation revolution that RL turned to be neural network based, with many gradient based convergence algorithms. Modern RL is often separated into two flavors, being model “free” and model “based” RL, if you’re here for that, scroll to the end.
## Markov Decision Processes (MDPs)
We start here because we need to learn about the model that is used to describe the world in most reinforcement learning problems.
### Why should I care?
Anyone interested in the growth of reinforcement learning should know the model they’re built on — Markov Decision Processes. They set up the structure of a world with **uncertainty** in where actions will take you, and agents need to **learn how to act**.
#### Non-Deterministic Search
Search is a central problem to artificial intelligence and intelligent agents. By planning into the future, search allows agents to solve games and logistical problems — but they rely on knowing where a certain action will take you. In traditional, **tree-based methods**, an action takes you to a next state, there is no distribution of next-states. That means, if you have the storage for it, you can plan **set, deterministic trajectories** into the future. Markov Decision Processes make this planning **stochastic,** or non-deterministic. The list of topics in search related to this article is long — [graph search](https://en.wikipedia.org/wiki/Graph_traversal), [game trees](https://en.wikipedia.org/wiki/Game_tree), [alpha-beta pruning](https://en.wikipedia.org/wiki/Alpha%E2%80%93beta_pruning), [minimax search](https://en.wikipedia.org/wiki/Minimax), [expectimax search](https://en.wikipedia.org/wiki/Expectiminimax), etc.
In the real world, this is a far better model for how agents act. Every simple action we take — pouring the coffee, sending a letter, moving a joint — has an expected outcome, but there is a sort of **randomness to life**. Markov Decision Processes are the tool that makes planning capture this uncertainty.
### What is Markov about MDPs?
Markov is all about Andrey Markov — a famous Russian mathematician most known for his work on stochastic processes.
“Markov” generally means that given the present state, the future and the past are independent.
#### [Andrey Markov (1856–1922)](https://en.wikipedia.org/wiki/Andrey_Markov)
The key idea of making a *Markovian* system is **memorylessness**. Memorylessness is the idea that the history of a system does not impact the current state. In probability notation, **memorylessness** translates into this. Consider a sequence of actions yields a trajectory, and we are looking to see where the current action will take us. The long conditional probability could look like:
Now — if the system is Markovian, the history is ***all encompassed in the current state***. So, our one step distribution is far simpler. This one step is a game-changer for computational efficiency. *The Markov Property underpins the existence and success of all modern reinforcement learning algorithms.*
### Markov Decision Process (MDPs)
An MDP is defined by the following quantities:
- Set of states **s ∈ S**. The states represent all the possible configurations of the world. In the example below, it is robot locations.
- Set of actions **a ∈ A**. The actions are the collection of all possible motions an agent can take. Below the actions are North, East, South, West.
- A transition function **T(s,a,s’)**. T(s,a,s’) holds the **uncertainty** of an MDP. Given a current position, and a provided action, T governs how frequently a certain next state follows. In the example below, the transition function could be that the next state is in the direction of the action 80% of the time, but is off by 90degrees the other 20%. How does this affect planning? In the example below, the robot chose North, but there’s a 10% chance each of it going East or West.
- A reward function **R(s,a,s’). Maximizing the sum of rewards is the goal of any agent.** This function says how much reward is gained at each step. In general, there will be a small negative reward (cost) at each step to encourage fast solutions, and big positive (goal) or negative (failed task) rewards at terminal states. Below, the Gem and the Fire Pit are the terminal states.
- A start state **s0**, and maybe a terminal state.
Example MDP from teaching [CS188 at UC Berkeley](https://miro.medium.com/max/2116/0*sB0cYflezQERRESh.png).
#### What does this give us?
This definition gives us a finite world, we a set forward dynamics model. We know the exact probabilities for each transition, and how good each action is. Ultimately, this model is a **scenario** — a scenario where we will plan how to act knowing that our actions may go slightly awry.
If the robot is next to the fire pit, should the robot always choose North knowing that there’s a chance that North will send it East?
No — the optimal policy will be West. Going into the wall will eventually (20% chance) go North, and put the robot on track for the goal.
### Policies
Learning how to act in an unknown environment is the final goal of understanding an environment. In MDPs, this is called the **policy**.
A policy is a function that gives you an action from a state. π*: S → A.
There are many methods of getting to a policy, but the core ideas are value and policy iteration. Both of these methods **iteratively build estimates for the total utility** of a state, and maybe an action.
The Utility of a state is the sum of (discounted) rewards.
Once every state has a utility, planning and policy generation at a high level **becomes following the line of maximum utility**. In MDPs and other learning approaches, the models add a **discount factor** γ to prioritize near term to long term rewards. The discount factor makes sense intuitively — humans and biological creates value money (or food) in hand now more than later. The discount factor also carries with it immense computational convergence help by changing the sum’s of rewards into a geometric series.
Photo break from [Pexels ](https://www.pexels.com/photo/aerial-photography-of-concrete-road-1646164/).
## The hidden linear algebra of RL
How do fundamentals of linear algebra support what is going on in MDPs. This section is something you won’t get in a normal intro course, but I find it fun to consider the iterative updates when solving Markov Decision Process as a dynamical system.
### Intro
How do fundamentals of linear algebra support the pinnacles of deep reinforcement learning? The answer is in the iterative updates when solving Markov Decision Process.
Reinforcement learning (RL) is the set of intelligent methods for **iteratively** learning a set of tasks. As computer science is a **computational** field, this learning takes place on vectors of states, actions, etc. and on matrices of dynamics or transitions. The states and vectors can take different forms, but how can we look at the convergence of the algorithms making headlines around the technology community? When we think of passing a vector of variables through some linear system, and getting a similar output, **eigenvalues** should come to mind.
$$ \large{ \lambda \overrightarrow{u}' = A \overrightarrow{u} }$$
#### *Important* values
There are two important characteristic utilities of a MDP — values of a state, and q-values of a chance node.
- **Value** of a state: The value of a state is the optimal recursive sum of rewards starting from a state. *The value of the state to the left would be greatly different if the robot is in a fire pit, near a gem, or on a couch.*
- **Q-value** of a state, action pair: The q-value is the optimal sum of discounted rewards associated with a state-action pair. *The q-value of a state is determined by an action — so the q-values on the ledge will vary greatly if pointing in or out of flames!*
These two values are related by [mutual recursion](https://en.wikipedia.org/wiki/Mutual_recursion), and Bellman updates.
### Bellman Updates
Richard E. Bellman was a mathematician that laid the groundwork for modern control and optimization theory. Through a recursive one-step equation, **a Bellman Update Equation,** large optimization problems can be solved efficiently. With a recursive Bellman update, one can set up an optimization or control problem with **Dynamic Programming,** which is a process of creating smaller, more computationally tractable problems. This process proceeds *recursively* from the end — a receding horizon approach.
Richard E. Bellman (1920–1984) — Wikipedia.
- **Bellman Equation**: *Necessary condition for optimality in optimization problems formulated as Dynamic Programming****.***
- **Dynamic Programing**: *Process to simplify an optimization problem by breaking it down into an optimal substructure.*
In reinforcement learning, we use the Bellman Update process to solve for the optimal values and q-values of a state-action space. This is ultimately formulating the **expected sum of future rewards** from a given location.
Here, we can see all of the values from the review interleaving. The notation **(*)**denotes optimal, so true or converged. We have the value of the state being determined by the best action, and a q-state, then two recursive definitions. The recursive values balance the *probability of visiting any state in* ***T(s,a,s’)*** and the *reward of any transition* ***R(s,a,s’)*** to create a global map for values of the state-action space*.*
The best value is tied to the best action-conditions q-value. Then the value and q-value update rules are very similar (weighting transitions, rewards, and discount). top) coupling of values to q-values; mid) Q-value recursion, bot) Value iteration. Source [cs188](https://inst.eecs.berkeley.edu/cs188/sp20/) at UC Berkeley.
They key point here is that we are multiplying matrices (***R, T***), by vectors (***V,U***), to iteratively solve for convergence. *The values will converge from any initial state because of how the values for one state are determined by their neighbors* ***s’***.
#### Reinforcement Learning?
“I was told there would be RL,” — you, reader, 8 minutes in. This is all reinforcement learning, and I assert **understanding the assumptions and the model that the algorithms are built on will prepare you vastly better** then just copying python tutorials from OpenAI. Do that after. I’ve mentored multiple students into working in RL, and the *ones who get more done are always the ones that learn what is going on, and then how to apply it.*
That being said, this is **one small step away from online q-learning**, where we estimate the same Bellman updates with samples of T and R rather than explicitly using them in the equations. **All the same assertions apply**, but it is over probability distributions and expectations. [Q-learning is the famous algorithm that solved Atari games and more in 2015](https://daiwk.github.io/assets/dqn.pdf).
### Hidden Math
#### Eigenvalues? Huh.
Recall an eigenvalue-eigenvector pair (*λ, u*) of a system *A* is a vector and scalar such that the vector acted on by the system returns a scalar multiple of the original vector.
The eigenvalue, eigenvector equation.
The beautiful thing about eigenvalues and eigenvectors is that when they span the state space (which they are guaranteed to do for most physical systems by something called generalized eigenvectors), every vector can be written as a combination of the other eigenvectors. Then, in discrete systems Eigenvectors control the evolutions from any initial state — any initial vector will combine to a **linear combination of the eigenvectors.**
#### Stochastic Matrices and Markov Chains
MDPs are very close to, but not the same in structure to Markov Chains. Markov chains are determined by transition matrix ***P***. The probability matrix acts like the transition matrix ***T(s,a,s’)*** *summed over the actions*. In Markov Chains, the next state is determined by:
$$ \large{ \overrightarrow{x}' = P \overrightarrow{x} }$$
Evolution of a [stochastic matrix.](https://en.wikipedia.org/wiki/Stochastic_matrix)
This matrix **P** has some special values — you can see that this is an *eigenvalue equation with all the eigenvalues equal to one* (picture a *λ*=1 pre-multiplying the left side of the equation). In order to get a matrix **guaranteed** to have eigenvalues equal to one, all the columns must sum up to 1.
What we are looking for in RL now, is how does the evolution of our solutions relate to convergence of probability distributions? We do this by formulating the iterative operators for ***V**** and ***Q**** as a linear operator (a Matrix) ***B***. Convergence can be tricky — the value and q-value vectors we use are not the eigenvectors — they converge to the eigenvectors, but that’s not important to seeing how **eigenvectors govern the system**.
$$ \large{ \overrightarrow{u}' = B \overrightarrow{u} }$$
The Bellman operator, B, like a linear transformation with eigenvector, eigenvalue pair of *λ=1.*
*Any initial value distribution will converge to the shape of the eigenspace. This illustration doesn’t show the exact eigenvalues of the Bellman update, but how the shape of the space could evolve as the values recursively update. Initially, the values will be totally unknown, but as learning emerges — the known values will converge to match the environment exactly.*
### Bellman Updates to Matrices
So far, we know that if we can formulate our Bellman updates (which are linear updates) in a simpler form, a convenient eigen-structure will emerge. How can we formulate our ***Q*** update as a simple update equation? We start with a Q-iteration equation (Optimal value substituted for the equivalent Q-value equation on the right).
Moving our system towards a linear operator (Matrix)
#### i) Let us rephrase a couple of the terms to **general forms**
The first half of the update, the summation over ***R, T,*** is an explicit reward number; we call it **R(s)**. Next, we change our summation over transitions to a probability matrix (matching a Markov matrix, convenient). Also, this leads into the next step — changing from ***Q*** to utilities. (Recall, you can take a maximization out of a sum (think of it as a more general upper bound).
Close to the Bellman matrix we want, where the P(s’|s,a) would dictate the evolution of our matrix.
#### ii) **Let’s make this a vector equation**
We are most interested in how the utility, ***U***, evolves for a MDP. That utility implies the value or q-value. We simply can re-write our ***Q*** into a ***U*** without much of a change, but that means we are assuming a fixed policy.
Transforming the q-state to a general utility vector for eigenvalues.
It’s important to recall that even for a multi-dimensional, physical system —*the utilities of a state are a vector if we stack all of the measured states into a long array*. A fixed policy doesn’t change convergence — it just means we have to revisit this to learn how to get a policy iteratively.
#### iii) Assume a **fixed policy**
If you assume a fixed policy, the maximization over ***a*** disappears. The maximization operator is distinctly nonlinear, but there are forms in linear algebra that are eigenvectors plus an additional vector (hint — [generalized eigenvectors](https://en.wikipedia.org/wiki/Generalized_eigenvector)).
This equation above is a general form of a Bellman update on Utility. We wanted a linear operator, ***B***, then we could see how this is an eigenvalue evolution equation. It looks a little different, but this is ultimately the form we want, minus a couple linear algebra assertions, so we have our Bellman update.
$$ \large{ \vec{u}'=B\vec{u} }$$
The Bellman operator, B, like a linear transformation with eigenvector, eigenvalue pair of *λ=1.*
Computationally, one can get the eigenvectors we want, but doing so analytically is challenging because of the assumptions made along the way.
### Takeaway
Linear operators show you how certain discrete, linear systems will evolve — and the environments we use in reinforcement learning follow that structure.
Eigenvalues and Eigenvectors of our collected data can represent the latent value space of a RL problem.
The nitty gritty of change of variables, linear transformations, fitting this into online q-learning (rather than the q-iteration here), and more will be in a future post.
## Fundamental Methods of RL & MDPs
This section focuses on taking an understanding of basic MDPs and applying it to how it relates to a fundamental reinforcement learning method. The methods I will focus on are **Value Iteration** and **Policy Iteration**. These two methods underpin **Q-value Iteration**, which directly leads to **Q-Learning**.
Q-Learning kick-started the deep reinforcement learning wave we are on, so it is a crucial peg in the reinforcement learning student’s playbook.
### Review Markov Decision Processes
Markov Decision Processes (MDPs) are the stochastic model underpinning reinforcement learning (RL). If you’re familiar, you can skip this section, but I added explanations for why each element matters in a *reinforcement learning context*.
***Revisiting Definitions & RL***
- Set of states **s ∈ S**, actions **a ∈ A**. The states and actions are the collection of all possible positions and motions of an agent. ***In advanced reinforcement learning****, the states and actions often become continuous, which requires a rethink of our algorithms.*
- A transition function **T(s,a,s’)**. Given a current position, and a provided action, ***T*** governs how frequently a certain next state follows. ***In reinforcement learning****, we no longer have access to this function, so the methods attempt to approximate it or learn implicit on sampled data.*
- A reward function **R(s,a,s’).** This function says how much reward is gained at each step. ***In reinforcement learning****, we no longer have access to this function, so we learn from sampled values* ***r*** *that lead the algorithms to explore environments and then exploit optimal trajectories.*
- A discount factor **γ (gamma)** in [0,1] which tunes the value of immediate (next step) to future rewards. ***In reinforcement learning****, we no longer have access to this function,* ***γ (gamma)*** *controls the convergence of most all learning algorithms and planning-optimizers through Bellman-like updates.*
- A start state **s0**, and maybe a terminal state.
### Leading towards reinforcement learning
#### Value Iteration
Learn the values for all states, then we can act according to the gradient. Value iteration learns the value of the states from the Bellman Update directly. The Bellman Update is guaranteed to converge to optimal values, under some non-restrictive conditions.
**Learning a policy may be more direct than learning a value**. Learning a value may take an infinite amount of time to converge to numerical precision of a 64bit float (think about a moving average averaging in a constant at every iteration, after starting with an estimate of 0, it will add a smaller and smaller nonzero number forever).
#### Policy Iteration
Learn a policy in tandem to the values. Policy learning incrementally looks at the current values and extracts a policy. Because the **action space is finite**, the hope is that it can converge faster than Value Iteration. Conceptually, the last change to the actions will happen well before the small rolling-average updates end. There are two steps to Policy Iteration.
The first is called **Policy Extraction**, which is how you go from a value to a policy — by taking the policy that maximizes over expected values.
The second step is **Policy Evaluation**. Policy evaluation takes a policy and runs value iteration *conditioned on a policy*. The samples are forever tied to the policy, but we know *we have to run the iterative algorithms for way fewer steps to extract the relevant* ***action*** *information*.
Policy evaluation step.
Like value iteration, policy iteration is guaranteed to converge for most reasonable MDPs because of the underlying Bellman Update.
#### Q-value Iteration
The problem with knowing optimal values is that it can be hard to distill a policy from it. The argmax operator is distinctly nonlinear and difficult to optimize over, so Q-value Iteration takes a step towards **direct policy extraction**. The optimal policy at each state is simply the max q-value at that state.
The reason most instruction starts with Value Iteration is that it slots into the Bellman updates a little more naturally. **Q-value Iteration requires the substitution of two of the key MDP value relations together**. After doing so, it is one step removed from Q-learning, which we will get to know.
### Iterative algorithms?
Let’s make sure you understand all the terms. Essentially, each update is made up of two terms after the summation (and potentially a max term that selects an action). Let’s factor out the brackets and discuss how they relate to an MDP.
$$ \large{ \sum_{s'} T(s,a,s') R(s,a,s')}$$
The first term is a summation over the product *T(s,a,s’)R(s,a,s’).* This term represents the latent value and likelihood of a given state and transition. The ***T*** term, or transition, governs how likely it is to get a given reward from a transition (recall, ***a*** tuple ***s,a,s’***determines a tuple where an action ***a*** takes an agent from state ***s*** to state ***s’***). This will do things like weight low probability states with high rewards against frequent states with lower rewards.
$$ \large{ \gamma \sum_{s'} T(s,a,s') V(s')}$$
The next term governs the **“Bellman-ness”** of these algorithms. It is a weighting of the data at the last step of the iterative algorithm — ***V***, with the term above. This pulls information from neighboring states about value so that we can understand longer-term transitions. *Think of this term as to where most of the recursive update happens, and the first term is a weighing prior determined by the environment.*
#### Conditions on convergence
All of the iterative algorithms are told to “converge to the optimal value or policy, under some conditions.” What are those conditions you ask?
- **Total state-space coverage**. The condition is that all state, action, next_state tuples are reached under the conditioned policy. Without this, some information from the MDP will be lost and Values can be stuck at the initial value.
- **Discount factor γ <1.** This is because the Values for any loop that can be repeated can and will go to infinity.
Thankfully, in practice, **these conditions are easy to meet**. Most exploration has an epsilon-greediness that includes a chance at a random action, always (so any action is feasible) and a non-one discount factor results in more favorable performance. Ultimately, these algorithms work in plenty of settings, so they’re definitely worth giving a shot.
### Reinforcement Learning
How do we make what we have seen into a reinforcement learning problem? We need to use samples rather than the true T(s,a,s’) and R(s,a,s’) functions.
#### Sample-based learning — how to solve a hidden MDP
The only difference between iterative methods in MDPs and the basic methods of solving a reinforcement learning problem is that **RL samples from the underlying transition and reward functions of an MDP**, rather than having it in the update rule. There are two things we need to update to get going, a replacement for ***T(s,a,s’)***and a replacement for ***R(s,a,s’)***.
First, let us approximate the transition function as the **average action conditioned transition for each observed tuple**. All the values we have not seen are initialized with random values. This is the simplest form of **model-based reinforcement learning** (my research area).
(Above:*An approximation for the transition function. The groundwork for advanced model-based reinforcement learning research.*)
Now, all that is left is remembering what to do with the reward, right? But, we actually have a reward with each step, so we can get away with it (methods average out to the correct value with many samples). Consider approximating the Q-value iteration equation with a sampled reward, as below.
Sample-based Q-learning (actual RL).
#### End note on Q-learning
**The above equation is Q-learning**. We start with some vector Q(s,a) that is filled with random values, and then we collect interactions with the world and tune alpha. Alpha is a learning rate, so we will lower it when we think our algorithm is converging.
It works out that Q-learning converges really similarly to Q-value Iteration, but we are just running the algorithm with an incomplete view of the world.
*The Q-learning used in robotics and games is in more complex feature spaces with neural networks approximating a large table of all the state-action pairs.* For a summary of how Deep Q-Learning shocked the world, [here is a great video](https://youtu.be/Ih8EfvOzBOY) until I write my own piece about it! But, how do these converge?
## Convergence of Iterative Algorithms
Are there any simple bounds one can put on the rate of learning a task? A study in the context of Q-learning.
Deep reinforcement learning algorithms may be the most difficult algorithms in recent machine learning developments to put **numerical bounds** on their performance (among those that function). The reasoning is twofold:
- Deep **neural networks are nebulous black boxes**, and no one truly understands how or why they converge so well.
- Reinforcement learning task **convergence is** **historically unstable** because of the sparse reward observed from the environment (and the difficulty of the underlying task — ***learn from scratch!***).
Here, I will walk you through a heuristic we can use to describe how RL algorithms can converge, and explain how to generalize it to more scenarios.
### Generalizing the iterative updates
In the first section of new content I will recall the RL concepts I am using, and highlight the **mathematical transformations** needed to get a system of equations that **evolves in discrete steps** and **has a convergence bound**. This section follows closely from the linear algebra section earlier.
#### Reinforcement Learning
RL is the paradigm where we are trying to “solve” and MDP, but we don’t know the underlying environment. The simple RL solutions are sampling-based variants of fundamental MDP-solving algorithms (Value and Policy Iteration). Recall Q-value Iteration, which is the Bellman Update I will focus on:
Looking at how accurate Value Iteration or Policy Iteration distills to comparing a value vector after each assignment (←) in the above equation, which is one round of the recursive update. The convergence of these methods yields a measure proportional to how reinforcement learning algorithms will converge because **reinforcement learning algorithms are sampling-based versions of Value and Policy Iteration, with a few more moving parts**.
*Recall: Q-learning is the same update rule as Q-value Iteration, but the transition function is replaced by the action of sampling and the reward function is replaced with the actual sample,* ***r****, received from the environment.*
#### Linear Operator and Eigenvectors
We need to formulate our Bellman Updates as a linear operator, ***B***,(a matrix is a subset of linear operators) and see if we can get it to behave as a [**stochastic matrix**](https://en.wikipedia.org/wiki/Stochastic_matrix). A stochastic matrix is [guaranteed](https://math.stackexchange.com/questions/40320/proof-that-the-largest-eigenvalue-of-a-stochastic-matrix-is-1) to have an **eigenvector** paired with the **eigenvalue 1**.
$$ \large{ \overrightarrow{u}' = B \overrightarrow{u} }$$
This says, we can study the iterative updates in RL like we would the evolution of an eigenspace where lambda is 1.
For now, we need to make a couple of notational changes to transition from formulas on *Q-values in a matrix* to formulas that are acting on **Utilities in a vector** (Utilities generalize to values and Q-values, anyways because it is defined as the discounted sum of rewards).
- We know utilities are proportional to q-values; **change the notation**. We will use ***U(s)***.
- Rearranging the utilities as a **vector** is like the flatten() function of many coding libraries (e.g. shifting an XY state-space to a 1d vector indexing across the states). We get ***u***
Transforming the q-state to a general utility vector for eigenvalues. These changes are crucial for formulating the problem as eigenvectors.
#### Understanding the Pieces of Q-learning
Start with the base equation for Q-value iteration below, how can we generalize this to a linear system? Recall the assumptions we made in the linear algebra section: 1) that we can average over the transition and reward distributions to make that a static vector ***R(s)*** and 2) the term with the Q values and the transition probabilities is like a ***stochastic matrix***.
This leaves us with a final form (merging equations from this section, and the end of the eigenvalue section). Can we use this as the linear operator that we need? Consider if we are using the optimal policy (without loss of generality), then the tricky maximum over actions drops out, leaving us with:
And, re-writing like a *generalized eigenvector*, we get:
$$ \large{\vec{u}=B\vec{u}+\vec{r} }$$
This is as closer as we can get to making this a pure eigenvector equation. It’s very close to what is called a [generalized eigenvector](https://en.wikipedia.org/wiki/Generalized_eigenvector).
### Studying Convergence
In this section, I will derive a relationship that guarantees a **minimum error of epsilon after N** steps and show what it means.
#### The system we can study
We saw above that we can formulate the utility update rule in a way that is very close to an eigenvector, but we were off by a constant vector ***r*** representing the underlying reward surface of an MDP. What happens if we take the difference between two utility vectors? The constant term drops out.
*Taking the difference between two utility vectors, we see that the ****r vector**** can cancel out.I*
We now have our system that we can study like a dynamical system of eigenvectors, but we are working in the space of the difference between utilities, with a matrix ***B*** *—***the sum of weighted transition probabilities**. Because the utility vectors are defined by each other, we can magically rearrange this (substitute the recursive Bellman Equation, take the norm). *The error after a Bellman update is reduced by the discount factor.*
Studying the difference between any utility estimate is ingenious, because it shows a) how an estimate differs from the true or b) how the data from only the recursive update evolves (not the little vector ***r***).
#### An epsilon-N Relation
Any convergence proof will be looking for a relationship between the **error bound**, ε, and the **number of steps**, *N*,(iterations). This relationship will give us the chance to bound the performance with an **analytical equation**.
We want the bound of our Utility error at step N — b(N) — to be less than epsilon. The bound will be an analytical expression, and epsilon is a scalar representing the norm of the difference between the estimated and true utility vectors.
$$ \large{ \epsilon > b(N) }$$
In other words, we are looking for a bound for epsilon that this a function of *N.*
To start, we know (by definition of an MDP) that the reward at each step, ***r***, is bounded in the interval [-Rmax, +Rmax]. Then, looking at the definition of utility (discounted sum of reward) as a geometric series, we can bound the difference from any vector to the true vector. *(*[*Proof for geometric series convergence*](https://en.wikipedia.org/wiki/Geometric_series#Proof_of_convergence)*).*
*First, recall the definition of utility: The expected sum of discounted reward.*
$$ \large{ U(s) = \sum_{t=0}^\infty \gamma^t r}$$
Next*, our starting point for convergence. Consider the difference between the true utility U and the initial utility U(0) from the convergence of a *[*geometric series*](https://en.wikipedia.org/wiki/Geometric_series)*. *
The bound comes from the worst-case estimate — where the true reward at every step is +Rmax, but we initialize our estimate to -Rmax. Alas, we have an **initial bound on the error of our utilities!** Recall U0 is what we initialize the utility vector to, and then the index will increase each time we run a Bellman Update.
To the Bellman Updates — **how does this bound at initialization evolve with each step**? Above, we see that **the error is reduced by the discount factor at each step**(from how the sum in the recursive update is always prepended with a gamma). This evolves below into a series of decreasing errors with each iteration.
The convergence of the difference between the current Utility update (U_i) and the true value **U**.
All that is left is declaring the bound, ***epsilon***, in relation to the number of steps, ***N***.
#### The Outcome — Visualizing Convergence
On the right-hand side of the equation above, we have a bound on the accuracy of our utility estimate.
Logically — for any epsilon, we know that the error will be less than epsilon in N steps.
We can also change back and forth between epsilon and N with a mathematical trick — so if we know *how accurate we want the estimates to be, we can solve for how long to let our algorithm run*!
$$ \large{ N = \frac{ log\frac{2R_{max}}{\epsilon(1-\gamma)}}{log(\frac{1}{\epsilon})} }$$
By taking the logarithm of both sides we can solve for the number of iterations to reach said bound!
The bound we have found is the bound on the cumulative value error across the state-space (solid line below). What the astute reader will wonder is, **how does the policy error compare?** Conceptually, this means, *‘in how many states would the current policy differ from the optimal?’* Turns out, with some normalization of the values (so they are numerically similar in magnitude) the policy error converges faster and gets to zero more rapidly (no asymptote!).
The convergence of value iteration vs policy iteration in policy lost space. The discrete nature of policies makes policy iteration converge faster under many circumstances. This represents the advantage in some situations to use Policy Iteration over Value Iteration. Does it carry to more algorithms?
### Bounding Recent Deep RL Algorithms
Bounding deep RL algorithms is what everyone wants. We have seen impressive results in recent years where robots can [run](https://arxiv.org/pdf/1812.11103.pdf), [fold towels](https://bair.berkeley.edu/blog/2018/11/30/visual-rl/), and [play](https://www.aaai.org/ocs/index.php/AAAI/AAAI18/paper/viewFile/16669/16677)games. It would be fantastic if we have bounds on performance.
What we can do, is bound how our representation of the world will converge. We have shown that **the utility function will converge**. There are two lasting challenges:
- *we are not given the* ***reward function*** *for real-life tasks, we must design it.*
- running these iterative algorithms is currently **not saf**e. Robotic exploration involves a lot of force, interaction, and (realistically) damage.
We can study algorithm convergence, but the majority of engineering problems limiting adoption of deep RL for real world tasks are reward engineering and safe learning.
That is where I leave you — a call to action to help us engineer better systems, so we can show off more of the underlying mathematics governing it. There’s an explanation of recent algorithmic advances below, but the fundamental understanding of what RL is finishes here.
## Gists of Some RL Algorithms (Sp 2019)
I conclude with a resource for getting the gist of RL algorithms without needing to surf through piles of documentation or equations.
A resource for getting the gist of RL algorithms without needing to surf through piles of documentation — a resource for students and researchers, without a single formula. As a reinforcement learning (RL) researcher I often need to remind myself of the subtle differences between the algorithms. Here I want to create a list of algorithms and a sentence or two for each that distinguishes it from others in its sub area. I pair this with a brief historical introduction to the field.
### Model Free RL
*Model free RL directly generates a policy for an actor. I like to think of it as end-to-end learning of how to act, with all the environmental knowledge being embedded in this policy.*
#### Policy Gradients Algorithms
*Policy gradient algorithms modify an agent’s policy to track those actions that bring it higher reward*. This lends these algorithms to be on-policy, so they can only learn from actions taken within the algorithm.
**Simple Statistical Gradient-Following Algorithms for Connectionist Reinforcement Learning** (REINFORCE) — [1992](https://link.springer.com/article/10.1007/BF00992696): This paper kickstarted the policy gradient idea, suggesting the core idea of systematically increasing the likelihood of actions that yield high rewards.
#### Value Based Algorithms
*Value based algorithms modify an agent’s policy based on the perceived value of a given state.*This lends these algorithms to be off-policy because an agent can update its internal value structure of a state by reading the reward function from any policy.
**Q-Learning** — [1992](https://link.springer.com/article/10.1007/BF00992698): Q-learning is the classic value based method in modern RL, where the agent stores a perceived value for each action, state pair, which then informs the policy action.
**Deep Q-Network** (DQN) — [2015](https://arxiv.org/pdf/1509.06461.pdf): Deep Q-Learning simply applies a neural network to approximate the Q function for each action and state, which can save vast amounts of computational resources, and potentially expand to continuous time action spaces.
#### Actor-Critic Algorithms
*Actor-critic algorithms take policy based and value based methods together — by having separate network approximations for the value (critic) and actions (actor). These two networks work together to regularize each other and create, hopefully, more stable results.* Summary of Actor Critic Algorithms: [source](https://arxiv.org/pdf/1708.05866.pdf)— Arulkumaran et al. “A Brief Survey of Deep Reinforcement Learning.”
**Actor Critic Algorithms** — [2000](https://papers.nips.cc/paper/1786-actor-critic-algorithms.pdf): This paper introduced the idea of having two separate, but intertwined models for generating a control policy.
#### Moving on From the Basics
A decade later, we find ourselves in an explosion of deep RL algorithms. Note that in all the press you read, deep at the core is referring to methods using neural network approximations.
Policy gradient algorithms regularly suffer from noisy gradients. I talked about one change in the gradient calculations recently proposed [in another post](https://medium.com/@natolambert/deep-rl-case-study-policy-based-vs-model-conditioned-gradients-in-rl-4546434c84b0), and a bunch of the most recent ‘State of the Art’ algorithms at their time looked to address this, including TRPO and PPO.
**Trust Region Policy Optimization** (TRPO) — [2015](http://proceedings.mlr.press/v37/schulman15.pdf): Building on the actor critic approach, the authors of TRPO looked to regularize the change in policies at each training iteration, and they introduce a hard constraint on the KL divergence[1](https://natolambert.me/writing/2020/0826_rl.html#fn1), or the information change in the new policy distribution. The use of a constraint, rather than a penalty, allows bigger training steps and faster convergence in practice.
**Proximal Policy Optimization** (PPO) — [2017](https://arxiv.org/abs/1707.06347): PPO builds on a similar idea as TRPO with the KL Divergence, and addresses the difficulty of implementing TRPO (which involves conjugate gradients to estimate the Fisher Information matrix), by using a surrogate loss function taking into account the KL divergence. PPO uses clipping to make this surrogate loss and assist convergence.
**Deep Deterministic Policy Gradient** (DDPG) — [2016](https://arxiv.org/pdf/1509.02971.pdf): DDPG combines improvements in Q learning with a policy gradient update rule, which allowed application of Q learning to many continuous control environments.
**Combining Improvements in Deep RL** (Rainbow) — [2017](https://arxiv.org/abs/1710.02298): Rainbow combines and compares many innovations in improving deep Q learning (DQN). There are many papers referenced here, so it can be a great place to learn about progress on DQN:
- Prioritization DQN: Replay transitions in Q learning where there is more uncertainty, ie more to learn.
- Dueling DQN: Separately estimates state values and action advantages to help generalize actions.
- A3C: Learns from multi-step bootstraps to propagate new knowledge to earlier in the networks.
- Distributional DQN: Learns a distribution of rewards rather than just the means.
- Noisy DQN: Employs stochastic layers for exploration, which makes action choices less exploitative.
The next two incorporate similar changes to the actor critic algorithms. Note that SAC is not a successor to TD3 as they were released nearly concurrently, but SAC uses a few of the tricks also used in TD3.
**Twin Delayed Deep Deterministic Policy Gradient** (TD3) — [2018](https://arxiv.org/abs/1802.09477): TD3 builds on DDPG with 3 key changes: 1) “Twin”: learns two Q functions simultaneously, taking the lower value for the Bellman estimate to reduce variance, 2) “Delayed”: updates the policy less frequently than the Q function, 3) adds noise to the to the target action to lower exploitative policies.
**Soft Actor Critic** (SAC) — [2018](https://arxiv.org/abs/1801.01290): To use model-free RL in robotic *experiment* the authors looked to improve sample efficiency, the breadth of data collection, and safety of exploration. Using entropy based RL they control exploration along with DDPG style Q function approximation for continuous control. *Note:* SAC also implemented clipping like TD3, and using a stochastic policy it benefits from regularizing action choice, which is similar to smoothing.
Many people are very excited about the applications of model-free RL as sample complexity falls and results rise. Recent research has brought an increasing portion of these methods to physical experiments, which is bringing the prospects of widely available robots one step closer.
### Model Based RL
*Model based RL(MBRL) attempts to build knowledge of the environment, and leverages said knowledge to take an informed action. The goal of these methods is often to reduce sample complexity on the model-free variants that are closer to end-to-end learning.*
**Probabilistic Inference for Learning Control** (PILCO) — [2011](https://www.ias.informatik.tu-darmstadt.de/uploads/Publications/Deisenroth_ICML_2011.pdf): This paper is one of the first in model-based RL, and it proposed a policy search method (essentially policy iteration) on top of a Gaussian Process (GP) dynamics model (built in uncertainty estimates). There have been many applications of learning with GPs, but not as many core algorithms to date.
**Probabilistic Ensembles with Trajectory Sampling** (PETS) — [2018](https://arxiv.org/abs/1805.12114): PETS combines three parts into one functional algorithm: 1) a dynamics model consisting of multiple randomly initialized neural networks (ensemble of models), 2) a particle based propagation algorithm, and 3) and simple model predictive controller. These three parts leverage deep learning of a dynamics model in a potentially generalizable fashion.
**Model-Based Meta-Policy-Optimization** (MB-MPO) — [2018](https://arxiv.org/abs/1809.05214): This paper uses meta-learning to choose which dynamics model in an ensemble best optimizes a policy and mitigate model bias. This meta-optimization allows MBRL to come closer to asymptotic model-free performance in substantially lower samples.
**Model-Ensemble Trust Region Policy Optimization** (ME-TRPO) — [2018](https://arxiv.org/abs/1802.10592): ME-TRPO is the application of TRPO on an ensemble of models assumed to be the ground truth of an environment. A subtle addition to the model-free version is a stop condition on policy training only when a user defined proportion of models in an ensemble no longer sees improvement when the policy is iterated.
**Model-Based Reinforcement Learning for Atari** (SimPLe) — [2019](https://arxiv.org/abs/1903.00374): SimPLe combines many tricks in the model-based RL area with a variational auto-encoder modeling dynamics from pixels. This shows the current state of the art for MBRL in Atari games (personally I think this is a very cool piece to read, and expect people to build on it soon).
The hype behind model-based RL has been increasing in recent years. It has often been given a short look because it lacks the asymptotic performance of its model-free counterparts. I am particularly interested in it because it has enabled many experiment only, exciting applications including: [quadrotors](https://arxiv.org/abs/1901.03737) and [walking robots](https://arxiv.org/abs/1708.02596).
### References and Resources:
A [great review of deep RL as of 2017](https://arxiv.org/pdf/1708.05866.pdf). And, these two:
#### [CS 294-112 Berkeley Deep RL Course (now of a new name CS285…)](http://rail.eecs.berkeley.edu/deeprlcourse/)
#### [Spinning Up in Deep RL](https://blog.openai.com/spinning-up-in-deep-rl/)
### Medium tries to save its writers
URL: https://natolambert.com/writing/medium
Date: 2021-02-16
Summary: Why Medium is not a website designed for the best writers.
The space of internet writing, paid subscriptions, and news distribution has been moving fast in 2020. With Substack on the rise (including multiple big name writers jumping ship for the direct-to-consumer model), Medium realized it needs to keep some of it’s writers from moving off platform.
*Medium for me is now a syndicate*. I write in my [own editor](https://ulysses.app/), [prioritize my direct writings](https://democraticrobots.substack.com/), and [host on my own website](https://natolambert.me/writing/index.html). Medium is the best route to traffic for now, but their paywall [really limits access](https://www.cake.co/conversations/06GgHsG/did-medium-just-put-all-of-their-content-behind-a-paywall). If you Google Medium Paywall Problem or something of the like, you get tons of titles and testimonials to how the paywall is limiting access (because they don’t distribute articles not in the paywall).
The average order of views I get with and without the paywall is striking:
- Without using their distribution (0-100s of views)
- Using their distribution, no publications (0-1,000s of views)
- Using a syndicate publication such as [towardsdatascience.com](https://towardsdatascience.com/@natolambert) (100s-10,000s of views)
The relationship between quality of what I write and these views is almost entirely separated, and that is the biggest problem for me. What I want: a platform with a clean reading interface where I know **my readers can follow me and connect with my content** — any risk of disruption to access is all the reason to leave, and disruption to followership is a barrier to enter.
*Here’s chronological events as Medium tries to save its writers.*
From Pexels, a Medium.
## [Medium Newsletters: A new way to connect with readers](https://blog.medium.com/medium-newsletters-4fe903b5bda7) (3 Jun 2020)
[This](https://blog.medium.com/medium-newsletters-4fe903b5bda7) is in response to [Substack](https://substack.com/), but it is too late. As a writer with thousands of views, “Top writer in AI” tag when I want it, hundreds of subscribers — I didn’t know about this when I made my newsletter.
### The problem: newsletters are only for publications
I **think** everyone can make their own publication on Medium, but this is limited because Medium stepped back on people’s abilities to have custom URLs. How do readers make the link it’s you if your email needs to go to the main medium.com URL.
A second problem: *what is the difference between ****following**** a writer and ****subscribing**** on medium?* Doing just one of these makes a website much more accessible. This also leads to: do I follow publications or authors? Keeping these trade-offs simple is what makes a publishing site run smoothly. As an author, how do I know my audience is following **me**.
This means, I can follow a writer and get more of their stories recommended to me, but is this enough? I have heard from multiple friends that following me on Medium by no means that they will see all of my content — and that is the specific reason they made an account. More on this trend is to follow: custom URLs, better author following, and more.
## [A more expressive Medium](https://blog.medium.com/a-more-expressive-medium-483e567b19ff) (28 Jul 2020)
This [post](https://blog.medium.com/a-more-expressive-medium-483e567b19ff) seems minor, but is a big direction of the trend: **they want to give writers their own space**.
We’re launching with a foundational set of controls around color, headers, type, and branding so that you can make a space on Medium that is uniquely yours. And this is just the beginning: we intend to evolve and build on these features over time, giving you even more flexibility to make Medium your own.
This is big because it is acknowledging the Medium itself is not the reason the reader is there. The author is. What this will do is move Medium further in the direction of an aggregator and further from a publication. Medium wants to host our content to take our content, but as a writer you should be sure to own what you write.
Paired with this is a new mobile app that brings **inline reading**, so there is less of a barrier to click an article — but only for publications. Medium is really prioritizing publications, so let’s see if they bring enough incentives to allow writers to run their own publications. If I post under my own name, are the features of the future going to be on their app?
And finally, **short form content**. Honestly, not a strong opinion here if they can minimize the clutter. I don’t think competing with Twitter has gone well, but maybe they can make a TikTok of writing — Medium does like to flex it’s algorithms.
### Our favorite Medium setting:
[Set the canonical link](https://help.medium.com/hc/en-us/articles/360033930293-Set-a-canonical-link) for all your posts. Have this be somewhere that treats writers well (e.g. your own website or Substack. What this does (in theory) is redirect the search engine optimization from traffic to where you want it.
## [A new Medium on mobile](https://blog.medium.com/a-new-medium-on-mobile-7ddeed24b231) (20 Aug 2020)
Authors are coming first [here](https://blog.medium.com/a-new-medium-on-mobile-7ddeed24b231) — literally, they’re now taking up the homepage instead of curate content. Now, Medium is [adding icons](https://miro.medium.com/max/1400/1*UwoWy82v4-oP-QF1C7XKag.png) from those who you read frequently on the home page, and only one or two articles. They call their new features ***shelves*** and it is tuned towards the individual. For authors, there’s things like easier writing and stats viewing, but really having a people-forward homepage is huge for a writer. Now I don’t have to worry about pleasing the algorithm when tuning my content to their platform (well, hopefully not as much, we will have to see if this is true).
## [What’s around the corner for Medium](https://blog.medium.com/whats-around-the-corner-for-medium-b79e8764c9cd) (30 Aug 2020)
[Custom domains, fairer distribution (claimed), better profiles](https://blog.medium.com/whats-around-the-corner-for-medium-b79e8764c9cd), and more is Medium’s core response to disruption. While Medium has seen a near exponential growth in viewership during 2020 (a crazy year for news), they don’t seem happy with how that is being converted into followership and engagement (I agree).
Following people over algorithms — this is a big change, and a step back from the first reason I left Medium.
The big change we announced with our new mobile app is that we are going to be putting more emphasis on following people and publications, over the algorithmic feed.
As for the custom domains, it’ll interesting to see if it is limited to xyz.medium.com or if it lets you link with hosting companies like Google Domains for true separation from the Google name. That would *maybe* be enough to bring me back. The biggest news is that curation will not be tied to monetization — ***I no longer have to choose between views and $$$ vs open access***, which is important to me.
All posts are now eligible for further distribution across Medium, whether they are paywalled or not
Thanks for reading, where do you write?
## Bookshelf Source
# My Complete Book List
With ratings of overall quality, how well it aligns to what I like, and how interesting the content alone is.
* 2026.05 *The Fall of Hyperion* by Dan Simmons
* 2026.04 *The Infinity Machine: Demis Hassabis, DeepMind, and the Quest for Superintelligence* by Sebastian Mallaby
* 2026.03 *Means of Ascent* by Robert A. Caro
* 2026.03 *Runnin' Down a Dream: How to Thrive in a Career You Actually Love* by Bill Gurley
* 2026.02 *The Path to Power* by Robert A. Caro
* 2026.01 *Hyperion* by Dan Simmons
* 2025.12 *Season of the Witch* by David Talbot
* 2025.11 *Too High and Too Steep: Reshaping Seattle’s Topography* by David Williams
* 2025.11 *Breakneck: China's Quest to Engineer the Future* by Dan Wang
* 2025.07 *Apple in China: The Capture of the World's Greatest Company* by Patrick McGee
* 2025.06 *The Optimist: Sam Altman, OpenAI, and the Race to Invent the Future* by Keach Hagey
* 2025.05 *Empire of AI: Dreams and Nightmares in Sam Altman’s OpenAI* by Karen Hao
* 2025.05 *Full Tilt: Ireland to India with a Bicycle* by Dervla Murphy
* 2025.05 *Dust (Silo #3)* by Hugh Howey
* 2025.04 *Shift (Silo #2)* by Hugh Howey
* 2025.04 *Careless People: A Cautionary Tale of Power, Greed, and Lost Idealism* by Sarah Wynn-Williams
* 2025.03 *Abundance* by Ezra Klein and Derek Thompson
* 2025.03 *Zero Days* by Ruth Ware
* 2025.02 *Travels with Charley: In Search of America* by John Steinbeck
* 2025.02 *The Creative Act: A Way of Being* by Rick Rubin
* 2025.02 *Boom: Bubbles and the End of Stagnation* by Byrne Hobart and Tobias Huber
* 2025.01 *A Song of Ice and Fire, Books 1 – 3* by George R. R. Martin
* 2025.01 *The Structure of Scientific Revolutions* by Thomas S. Kuhn
* 2024.12 *The Nvidia Way: Jensen Huang and the Making of a Tech Giant* by Tae Kim
* 2024.12 *Outlive: The Science & Art of Longevity* by Peter Attia with Bill Gifford
* 2024.08 *The Anxious Generation* by Jonathan Haidt: **4** / 4 / 4
* 2024.08 *The Demon of Unrest: A Saga of Hubris, Heartbreak, and Heroism at the Dawn of the Civil War* by Erik Larson: **4** / 4 / 4
* 2024.08 *Burn Book* by Kara Swisher: **5** / 5 / 5
* 2024.07 *The Road* by Cormac McCarthy: **4** / 3 / 4
* 2024.07 *The Women* by Kristin Hannah
* 2024.05 *Death's End* by Cixin Liu
* 2024.04 *The Dark Forest* by Cixin Liu
* 2024.06 *Killers of the Flower Moon* by David Grann: **5** / 4 / 5
* 2024.05 *Going Infinite* by Michael Lewis
* 2024.03 *Trust* by Hernan Diaz
* 2024.01 *American Gun* by Cameron McWhirter and Zusha Elinson
* 2024.01 *Steve Jobs* by Walter Isaacson
* 2024.01 *The Worlds I See* by Fei-Fei Li
* 2024.01 *Recursion* by Blake Crouch
* 2023.12 *A Gentleman in Moscow* by Amor Towles
* 2023.10 *The Storyteller* by Dave Grohl
* 2023.09. *American Prometheus* by Kai Bird and Martin Sherwin: **4** / 4 / 5
* 2023.09. *The Boys in the Boat* by Daniel James Brown: **4** / 5 / 5
* 2023.08. *Open* by Andre Agassi: **5** / 5 / 5
* 2023.07. *Once there were Wolves* by Charlotte McConaghy: **4** / 4 / 4
* 2023.06. *All the Dangerous Things* by Stacy Willingham: **3** / 3 / 4
* 2023.06. *The Signal and the Noise* by Nate Silver: **4** / 5 / 5
* 2023.04. *The Creative Act: A Way of Being* by Rick Rubin: **3** / 3 / 4
* 2023.03. *A Flicker in the Dark* by Stacy Willingham: **4** / 3 / 4
* 2023.03. *Dune Messiah* by Frank Herbert: **3** / 4 / 4
* 2023.03. *Two Wheels Good: The History and Mystery of the Bicycle* by Jody Rosen: **4** / 5 / 4
* 2023.02. *Travels with Charley: In Search of America* by John Steinbeck: **4** / 4 / 4
* 2023.01. *Fire & Blood* by George R. R. Martin: **4** / 4 / 4
* 2023.01. *The Dragons of Eden* by Carl Sagan: **5** / 5 / 5
* 2023.01. *The Nightingdale* by Kristin Hannah: **4** / 4 / 4
* 2023.01. *The Wise Heart: A Guide to the Universal Teachings of Buddhist Psychology* by Jack Kornfield: **4** / 4 / 4
* 2022.12. *The Great Alone* by Kristin Hannah: **5** / 4 / 5
* 2022.11. *The Lincoln Highway* by Amor Towles: **4** / 4 / 4
* 2022.1. *Consolations* by David Whyte: **5** / 4 / 5
* 2022.09. *Four Winds* by Kristin Hannah: **4** / 4 / 4
* 2022.09. *Slaughterhouse Five* by Kurt Vonnegut: **4** / 3 / 4
* 2022.07. *Do Androids Dream of Electric Sheep?* by Philip. K. Dick: **4** / 4 / 4
* 2022.07. *Finding the Mother Tree: Discovering the Wisdom of the Forest* by Suzanne Simard : **4** / 5 / 4
* 2022.07. *The Invention of Nature: Alexander Von Humboldt's New World* by Andrea Wulf: **5** / 5 / 5
* 2022.07. *The Subtle Art of Not Giving a F\*ck* by Mark Manson: **3** / 3 / 4
* 2022.05. *Ready Player One* by Ernest Cline: **2** / 4 / 2
* 2022.05. *Revolutionary Power: An Activist's Guide to the Energy Transition* by Shalanda Baker: **3** / 4 / 3
* 2022.05. *Snow Crash* by Neal Stephenson: **4** / 4 / 4
* 2022.05. *The Alignment Problem: Machine Learning and Human Values* by Brian Christian: **4** / 5 / 4
* 2022.04. *Harry Potter and the Deathly Hallows* by J. K. Rowling: **4** / 4 / 5
* 2022.04. *Project Hail Marry* by Andy Weir: **5** / 5 / 5
* 2022.04. *The Overstory* by Richard Powers: **5** / 5 / 5
* 2022.03. *Atomic Habits: An Easy & Proven Way to Build Good Habits & Break Bad Ones* by James Clear: **4** / 5 / 5
* 2022.03. *Endurance: Shackleton's Incredible Voyage* by Alfred Lansing: **5** / 5 / 5
* 2022.03. *Harry Potter and the Half Blood Prince* by J. K. Rowling: **4** / 3 / 4
* 2022.03. *Harry Potter and the Order of the Phoenix* by J. K. Rowling: **5** / 4 / 5
* 2022.02. *Harry Potter and the Goblet of Fire* by J. K. Rowling: **4** / 4 / 4
* 2022.02. *Lewis Hamilton: Five-Time World Champion: The Biography* by Frank Worrall: **3** / 3 / 4
* 2022.01. *Harry Potter and the Prisoner of Azkaban * by J. K. Rowling: **5** / 3 / 5
* 2022.01. *Termination Shock* by Neal Stephenson: **2** / 4 / 3
* 2021.12. *Bewilderment* by Richard Powers: **5** / 4 / 4
* 2021.12. *Harry Potter and the Chamber of Secrets* by J. K. Rowling: **3** / 3 / 3
* 2021.12. *The Source* by James A. Michener: **5** / 4 / 5
* 2021.11. *Harry Potter and the Sorcerer's Stone* by J. K. Rowling: **4** / 3 / 4
* 2021.11. *What Happened to You?: Conversations on Trauma, Resilience, and Healing* by Oprah Winfrey and Bruce D Perry: **5** / 4 / 5
* 2021.1. *Before the Coffee Gets Cold* by Toshikazu Kawaguchi: **3** / 2 / 3
* 2021.1. *Endure: Mind, Body, and the Curiously Elastic Limits of Human Performance* by Alex Hutchinson: **4** / 5 / 4
* 2021.1. *The Anthropocene Reviewed: Essays on a Human-Centered Planet* by John Green: **5** / 5 / 5
* 2021.09. *Atlas of AI: Power, Politics, and the Planetary Costs of Artificial Intelligence* by Kate Crawford: **5** / 5 / 5
* 2021.09. *Stealth War* by Robert Spalding: **3** / 3 / 4
* 2021.09. *The Friend* by Sigrid Nunez: **4** / 3 / 5
* 2021.08. *Scientific Models in Philosophy of Science* by Daniela M. Bailer-Jones: **5** / 5 / 5
* 2021.08. *The Wild Robot* by Peter Brown: **4** / 3 / 5
* 2021.07. *Infinite Powers: How Calculus Reveals the Secrets of the Universe* by Steven Strogatz: **5** / 5 / 5
* 2021.07. *The Bomber Mafia* by Malcolm Gladwell: **5** / 5 / 5
* 2021.07. *The Splendid and the Vile* by Erik Larson: **4** / 4 / 5
* 2021.07. *This Is Your Mind on Plants* by Michael Pollan: **4** / 4 / 3
* 2021.07. *Under a White Sky: The Nature of the Future* by Elizabeth Kolbert: **5** / 5 / 5
* 2021.06. *The Song of Achilles* by Madeline Miller: **3** / 2 / 3
* 2021.05. *Barbarian Days: A Surfing Life* by William Finnegan : **5** / 4 / 5
* 2021.05. *Braiding Sweetgrass* by Robin Wall Kimmerer: **3** / 3 / 4
* 2021.04. *Genius Makers: The Mavericks Who Brought AI to Google, Facebook, and the World* by Cade Metz: **5** / 5 / 4
* 2021.04. *"Surely You're Joking, Mr. Feynman!": Adventures of a Curious Character* by Richard P. Feynman: **5** / 5 / 5
* 2021.04. * The Rise and Fall of the Third Reich: A History of Nazi Germany* by William L. Shirer: **5** / 4 / 5
* 2021.03. *Alone on the Wall* by Alex Honnold: **3** / 4 / 5
* 2021.03. *How to Avoid a Climate Disaster* by Bill Gates: **3** / 5 / 4
* 2021.03. * Rossumovi Univerzální Roboti (R.U.R.)* by Karel Čapek: **5** / 5 / 5
* 2021.02. *Artemis* by Andy Weir: **4** / 3 / 4
* 2021.02. *Untamed* by Glennon Doyle: **5** / 5 / 5
* 2021.01. *Mythos* by Stephen Fry: **4** / 3 / 4
* 2021.01. *Through the Language Glass: Why the World Looks Different in Other Languages* by Guy Deutscher: **4** / 4 / 5
* 2021.01.*The Three-Body Problem* by Cixin Liu: **4** / 4 / 5
* 2020.12. *Breath: The New Science of a Lost Art* by James Nestor: **2** / 4 / 3
* 2020.12. *Greenlights* by Matthew McConaughey: **5** / 5 / 5
* 2020.12. *New Laws of Robotics: Defending Human Expertise in the Age of AI* by Frank Pasquale: **3** / 5 / 3
* 2020.11. *A Promised Land* by Barack Obama: **4** / 2 / 4
* 2020.08. *The Thrilling Adventures of Lovelace and Babbage: The (Mostly) True Story of the First Computer* by Sydney Padua: **4** / 3 / 4
* 2020.07. *Race After Technology* by Ruha Benjamin: **3** / 3 / 3
* 2020.06. *Human Compatible: Artificial Intelligence and the Problem of Control* by Stuart Russell: **4** / 5 / 5
* 2020.06. *The Pillars of the Earth* by Ken Follett: **5** / 5 / 5
* 2020.05. *The Gatekeepers: How the White House Chiefs of Staff Define Every Presidency* by Chris Whipple: **5** / 3 / 5
* 2020.05. *The Most Dangerous Branch: Inside the Supreme Court in the Age of Trump* by David A. Kaplan : **3** / 2 / 4
* 2020.05. *The Righteous Mind: Why Good People Are Divided by Politics and Religion* by Jonathan Haidt: **5** / 3 / 5
* 2020.03. *Conscious: A Brief Guide to the Fundamental Mystery of the Mind* by Annaka Harris: **4** / 5 / 3
* 2020.03. *Free Will* by Sam Harris: **3** / 3 / 4
* 2020.02. *Draft Animals: Living the Pro Cycling Dream (Once in a While)* by Phil Gaimon: **5** / 5 / 5
* 2020.01. *What I Talk about When I Talk about Running: A Memoir* by Haruki Murakami: **4** / 3 / 4
* 2020.00 *10% Happier: How I Tamed the Voice in My Head, Reduced Stress Without Losing My Edge, and Found Self-Help That Actually Works - A T* by Dan Harris: **4** / 4 / 4
* 2020.00 *Becoming* by Michelle Obama: **5** / 4 / 5
* 2020.00 *Bruce Lee: A Life* by Matthew Polly: **5** / 4 / 5
* 2020.00 *Caffeine: How Caffeine Created the Modern World* by Michael Pollan: **4** / 4 / 4
* 2020.00 *Elon Musk: Tesla, Spacex, and the Quest for a Fantastic Future* by Ashlee Vance: **4** / 5 / 4
* 2020.00 *How to Be an Antiracist* by Ibram X. Kendi: **4** / 3 / 4
* 2020.00 *How to Change Your Mind: What the New Science of Psychedelics Teaches Us about Consciousness, Dying, Addiction, Depression, and Transcendence* by Michael Pollan: **5** / 3 / 5
* 2020.00 *Rising out of Hatred: The Awakening of a Former White Nationalist* by Eli Saslow: **5** / 3 / 5
* 2020.00 *Superintelligence* by Nick Bostrom: **3** / 4 / 5
* 2020.00 *The Botany of Desire* by Michael Pollan: **5** / 5 / 5
* 2020.00 *The Lord of the Rings* by J. R. R. Tolkien: **5** / 5 / 5
* 2020.00 *The Lost City of Z: A Tale of Deadly Obsession in the Amazon* by David Grann: **4** / 4 / 5
* 2020.00 *The Omnivore's Dilemma* by Michael Pollen: **4** / 5 / 4
* 2020.00 *White Fragility: Why It's So Hard for White People to Talk About Racism* by Robin DiAngelo: **4** / 3 / 4
* 2019.00 * Barbarians at the Gate: The Fall of RJR Nabisco* by Bryan Burrough and John Helyar: **5** / 3 / 5
* 2019.00 *Godel, Escher, Bach: An Eternal Golden Braid* by Douglas R. Hofstadter : **5** / 5 / 5
* 2019.00 *Homo Deus: A Brief History of Tomorrow* by Yuval Noah Harari: **3** / 4 / 3
* 2019.00 *The Visual Display of Quantitative Information* by Edward Tufte: **5** / 5 / 5
* 2018.06. *Deep Nutrition: Why Your Genes Need Traditional Food* by Catherine Shanahan: **5** / 5 / 5
* 2018.04. *A Brief History of Time* by Stephen Hawking: **5** / 5 / 5
* 2018.00 *Mistakes Were Made (But Not by Me): Why We Justify Foolish Beliefs, Bad Decisions, and Hurtful Acts* by Carol Tavris & Elliot Aronson: **4** / 4 / 4
* 2018.00 *Society of Mind* by Marvin Minsky: **4** / 5 / 5
* 2018.00 *Why We Sleep: Unlocking the Power of Sleep and Dreams* by Matthew Walker: **5** / 5 / 5
* 2017.09. *Tribe: On Homecoming and Belonging* by Sebastian Junger: **5** / 5 / 5
* 2017.05. *Sapiens* by Yuval Noah Harari: **4** / 3 / 5
* 2017.00 *Beyond Training: Mastering Endurance, Health & Life* by Ben Greenfield: **4** / 5 / 4
* 2017.00 *Predictably Irrational: The Hidden Forces That Shape Our Decisions* by Dan Ariely: **3** / 5 / 4
* 2017.00 *Shoe Dog: A Memoir by the Creator of Nike* by Phil Knight: **4** / 4 / 4
* 2016.00 *Becoming a Supple Leopard* by Kelly Starrett: **5** / 5 / 5
* 2016.00 *Born to Run: A Hidden Tribe, Superathletes, and the Greatest Race the World Has Never Seen* by Christopher McDougall: **5** / 4 / 5
* 2016.00 *How to Win Friends and Influence People* by Dale Carnegie: **4** / 3 / 5
## Talks And Slides
- May 2026: Open-models: The lens in China, adoption trends, and new usage patterns (Center for Security and Emerging Technology (CSET)) - https://natolambert.com/slides/china-atom-2026/talk.html
- March 2026: An Introduction to Reinforcement Learning from Human Feedback and Post-training (SALA 2026, Quito, Ecuador) - https://rlhfbook.com/teach/SALA-2026/
- February 2026: Building Language Models in the Era of Agents (Amazon AI, Seattle WA)
- February 2026: Building OLMo in the Era of Agents (LTI Colloquium @ Carnegie Mellon University)
- December 2025: Building Olmo 3 Think (Foundations of Reasoning in Language Models @ NeurIPS) - https://www.youtube.com/watch?v=uaZ3yRdYg8A
- December 2025: Good researchers obsess over evals (Evaluating the Evolving LLM Lifecycle @ NeurIPS) - https://youtu.be/uaZ3yRdYg8A?si=31zxbDFqqqXHwJIR&t=2465
- October 2025: (Keynote) Olmo-Thinking: Training a Fully Open Reasoning Model (PyTorch Conference) - https://www.youtube.com/watch?v=uolCS_94c4A
- October 2025: Open Models in 2025 (The Curve) - https://youtu.be/FUcilE5Gx_0
- October 2025: Building a thinking Olmo (COLM ScalR Workshop)
- June 2025: The art of a good (reasoning) model (Enterprise AI Agent Summit, Seattle WA) - https://www.youtube.com/watch?v=VAzL8RHot1c
- June 2025: A taxonomy for next-generation reasoning models (AI Engineer World's Fair) - https://www.youtube.com/live/-9E9_21tx04?si=SJ3KETsRrNvPzeOM&t=11348
- April 2025: Experiments with Reinforcement Learning with Verifiable Rewards (USC Information Sciences Institute) - https://www.youtube.com/watch?v=zYeIqzULzr0
- March 2025: The RL Era of Language Models (UC Santa Cruz, Silicon Valley Extension) - https://www.youtube.com/watch?v=J1APR8Bo9dE
- February 2025: An unexpected RL renaissance (Minds and Machines Seminar, Seattle) - https://www.youtube.com/watch?v=YXTYbr3hiFU
- December 2024: (Tutorial) Language Modeling (Neural Information Processing Systems) - https://neurips.cc/virtual/2024/tutorial/99526
- December 2024: How to approach post-training for AI applications (Infer AI Engineer Vancouver) - https://www.youtube.com/watch?v=grpc-Wyy-Zg
- December 2024: The state of reasoning (Latent Space @ NeurIPS) - https://www.youtube.com/watch?v=2pHE9L4ZZXM
- November 2024: Tulu 3 Preview (Princeton AI Alignment and Safety Seminar) - https://www.youtube.com/watch?v=ltSzUIJ9m6s
- May 2024: Life after DPO (Stanford CS224N: NLP with Deep Learning) - https://www.youtube.com/watch?v=dnF463_Ar9I
- April 2024: Aligning Open Language Models (Stanford CS25: Transformers United V4) - https://www.youtube.com/watch?v=AdLgPmcrXwQ
- March 2024: RewardBench: Evaluating Reward Models for Language Modeling (Deep Learning Classics and Trends)
- December 2023: History and Risks of RLHF (Workshop on Sociotechnical AI Safety)
- November 2023: Bridging RLHF from LLMs back to control (CoRL LangRob Workshop) - https://www.youtube.com/live/ThgGAZF4hgI?si=EUGqtGr1XqIELhIj&t=8827
- August 2023: Objective Mismatch in Reinforcement Learning from Human Feedback (Deep Learning Classics and Trends) - https://drive.google.com/file/d/1F1lIi48PrWl5_80wQF40Y5dQmCBiZUUh/view?usp=sharing
- July 2023: (Tutorial) Reinforcement Learning from Human Feedback (International Conference on Machine Learning) - https://slideslive.com/39004357
- June 2023: (Tutorial) Steering language models with RLHF and Constitutional AI (ACM Conference on Fairness, Accountability, and Transparency)
- March 2023: Reinforcement Learning from Human Feedback: Open and Academic Progress (UCL Dark Lab) - https://www.youtube.com/watch?v=8SgKDSX-Me0
- July 2022: Reward Reports for Reinforcement Learning (ICML Workshop on Responsible Decision Making) - https://slideslive.com/38988545
- April 2022: Planning through Exploration and Exploitation in MBRL (University of Pennsylvania PAL Group) - https://youtu.be/HxbeQI7hfo4
- March 2022: (Job Talk) Legible Reinforcement Learning via Dynamics Models (Microsoft Research)
- February 2022: (Job Talk) Synergy of Prediction and Control in MBRL (Amazon Robotics & AI)
- March 2021: Improving Model Predictive Control in MBRL (Cornell Robotics Seminar) - https://www.youtube.com/watch?v=H5Q1UAEkhZY
- April 2020: Model Learning for Low-level Control in Robotics (UC Berkeley Semiautonomous Seminar) - https://youtu.be/Z4EpSWfU48E