Ryan Kidd

CEO & Co-Founder @ MATS Research
977 karmaJoined Working (0-5 years)Berkeley, CA, USA
ryankidd.ai

Bio

Participation
6

Personal website: ryankidd.ai

Give me feedback! :)

Comments
58

Thanks for publishing this, Arb! I have some thoughts, mostly pertaining to MATS:

  1. MATS believes a large part of our impact comes via accelerating researchers who might still enter AI safety, but would otherwise take significantly longer to spin up as competent researchers, rather than converting people into AIS researchers. MATS highly recommends that applicants have already completed AI Safety Fundamentals and most of our applicants come from personal recommendations or AISF alumni (though we are considering better targeted advertising to professional engineers and established academics). Here is a simplified model of the AI safety technical research pipeline as we see it.

    Why do we emphasize acceleration over conversion? Because we think that producing a researcher takes a long time (with a high drop-out rate), often requires apprenticeship (including illegible knowledge transfer) with a scarce group of mentors (with high barrier to entry), and benefits substantially from factors such as community support and curriculum. Additionally, MATS' acceptance rate is ~15% and many rejected applicants are very proficient researchers or engineers, including some with AI safety research experience, who can't find better options (e.g., independent research is worse for them). MATS scholars with prior AI safety research experience generally believe the program was significantly better than their counterfactual options, or was critical for finding collaborators or co-founders (alumni impact analysis forthcoming). So, the appropriate counterfactual for MATS and similar programs seems to be, "Junior researchers apply for funding and move to a research hub, hoping that a mentor responds to their emails, while orgs still struggle to scale even with extra cash."
  2. The "push vs. pull" model seems to neglect that e.g. many MATS scholars had highly paid roles in industry (or de facto offers given their qualifications) and chose to accept stipends at $30-50/h because working on AI safety is intrinsically a "pull" for a subset of talent and there were no better options. Additionally, MATS stipends are basically equivalent to LTFF funding; scholars are effectively self-employed as independent researchers, albeit with mentorship, operations, research management, and community support. Also, 63% of past MATS scholars have applied for funding immediately post-program as independent researchers for 4+ months as part of our extension program (many others go back to finish their PhDs or are hired) and 85% of those have been funded. I would guess that the median MATS scholar is slightly above the level of the median LTFF grantee from 2022 in terms of research impact, particularly given the boost they give to a mentor's research.
  3. Comparing the cost of funding marginal good independent researchers ($80k/year) to the cost of producing a good new researcher ($40k) seems like a false equivalence if you can't have one without the other. I believe the most taut constraint on producing more AIS researchers is generally training/mentorship, not money. Even wizard software engineers generally need an on-ramp for a field as pre-paradigmatic and illegible as AI safety. If all MATS' money instead went to the LTFF to support further independent researchers, I believe that substantially less impact would be generated. Many LTFF-funded researchers have enrolled in MATS! Caveat: you could probably hire e.g. Terry Tao for some amount of money, but this would likely be very large. Side note: independent researchers are likely cheaper than scholars in managed research programs or employees at AIS orgs because the latter two have overhead costs that benefit researcher output.
  4. Some of the researchers who passed through AISC later did MATS. Similarly, several researchers who did MLAB or REMIX later did MATS. It's often hard to appropriately attribute Shapley value to elements of the pipeline, so I recommend assessing orgs addressing different components of the pipeline by how well they achieve their role, and distributing funds between elements of the pipeline based on how much each is constraining the flow of new talent to later sections (anchored by elasticity to funding). For example, I believe that MATS and AISC should be assessed by their effectiveness (including cost, speedup, and mentor time) at converting "informed talent" (i.e., understands the scope of the problem) into "empowered talent" (i.e., can iterate on solutions and attract funding/get hired). This said, MATS aims to improve our advertising towards established academics and software engineers, which might bypass the pipeline in the diagram above. Side note: I believe that converting "unknown talent" into "informed talent" is generally much cheaper than converting "informed talent" into "empowered talent."
  5. Several MATS mentors (e.g., Neel Nanda) credit the program for helping them develop as research leads. Similarly, several MATS alumni have credited AISC (and SPAR) for helping them develop as research leads, similar to the way some Postdocs or PhDs take on supervisory roles on the way to Professorship. I believe the "carrying capacity" of the AI safety research field is largely bottlenecked on good research leads (i.e., who can scope and lead useful AIS research projects), especially given how many competent software engineers are flooding into AIS. It seems a mistake not to account for this source of impact in this review.

To be clear, I think that an independent analysis from a capable team is obviously to be preferred over an internal or grantmaker-led analysis. A GiveWell for AI safety would be great! However, my point was that this would likely require hiring serious AI safety strategy expertise (not just general data science and impact analysis experience), which might be quite hard in the current job market, where AI safety grantmakers are really struggling to hire. Why would AI safety be exceptional in this way? Because the bulk of downstream impact of AI safety interventions is rooted in the reduction of existential risk, which is unobservable directly in the near-term, and near-term proxies might be highly misleading (plenty to discuss here). None of this means that we shouldn't try! I endorse a "GiveWell for AI safety" and I would fund this, if not for the fact that there would be an obvious COI.

During my first month working in AI safety, someone told me, "When you enter AI safety, you have to work out for the first time what tier of god-tier you are." At the time, I interpreted this as "just because you have a physics PhD and led a successful EA group for three years, this doesn't mean you are qualified/capable to do AI safety research." Many times, working in this space, I have been forced to confront my limitations. I hope that AI safety can be/become a much broader church than the comment I heard implies. I do not believe that a PhD is necessary to do AI safety research. I do think the field is very small and growth is constrained by strong founders, field-builders, and grantmakers, which has resulted in a highly elite set of programs and job market. I hope we can increasingly move to a place of abundance and enable more people to contribute their skills and dedication.

What's your evidence that "cG regularly turns away competent, overqualified applicants"?

I think it's deeply unfortunate if a lot of people are getting turned off AI safety as a cause area because they can't break into the field. I also think the calls for urgency are justified on impact grounds. Maybe there are low-cost ways to help rejected applicants feel better? I like the idea of publishing application statistics or offering tailored advice to rejected applicants, though the latter is really expensive 

It seems like a lot of people are rejected by Harvard or Google, but are able to pivot to other Ivy League colleges or FAANG companies. At worst, there are second-tier colleges and companies to apply to. I'm not sure this is the case in AI safety, as the field is still very small (2-4k FTEs). It seems like the best thing many rejected applicants can do is apply for a CS PhD or work in a regular tech company to build their skills, which can be deeply unsatisfying for impact-driven people who think AGI is near (as I do).

Do you think the marketing is false because:

  • The roles are getting more than enough qualified, competent applicants?
  • The cause area of AI safety is saturated with talent?
  • AI safety is not impactful enough to justify claims like "this cause area is particularly important to scale"?

I'm not sure the analogies to Google or Harvard are useful here. These organizations are driven by profits/enrollments, not by impact. They likely are far less picky than AI safety orgs because they are very big and don't require the incredibly niche skill-sets and mission alignment of AI safety roles. Google can also pay a ton and both Harvard and Google are conventionally high status, attracting a ton of applicants naturally. One might argue that there is a moral imperative to scale AI safety fast in a way that doesn't apply to Harvard or Google on the whole. Note that the GDM interpretability team and Harvard AI safety team have called for urgency, independent of Google's general messaging.

None of this is to justify the hiring practices at AI safety organizations. I think plenty of operations roles (e.g., HR, finance, legal, marketing, facilities) don't require deep mission alignment or specialist AI safety skills. At MATS, the hiring rate for ops roles is ~2%; there are just a ton of applicants! Obviously, roles that depend on specialist AI safety skills, knowledge, or connections will be more selective, as well as leadership or independently operating roles, which greatly benefit from deep mission alignment.

In defence of the call for urgency in combination with the low acceptance rates:

  • AI safety seems really impactful and urgent;
  • Marginally better talent can have outsized impacts in a fast-growing field;
  • Some roles (e.g., grantmakers) are very hard to hire for as the necessary skills are very rare;
  • Emphasizing the important and urgency of AI safety will allow more samples from the distribution of talent, which means better-staffed roles and higher likelihood of finding rare talent.

Ah, makes sense. In the world where MATS randomizes over applicants for a typical 100+ fellow cohort and we learn after 6-18 months how impactful MATS is, what are the main upsides? My guess:

  • Might help understand how useful fellowship programs are vs. other AI safety interventions, like academic grants or mass-messaging (most of which are also saturated with funding at the moment).
  • Might help better price how cost-effective AI safety is vs. other EA cause areas.
  • Might help MATS staff (e.g., me) work on higher-impact interventions.

Valid points, all. Re. your last point, if one thinks the default case is doom, once might want to fund a lot of blue sky interventions as the upside potential is stronger than the downside risk.

Steelmanning the case for RCTs at MATS:

  • Maybe we are actively causing harm by selecting fellows suboptimally and could improve fast and cheaply.
  • Maybe we can show mentors convincing evidence that helps them select fellows better without risking their support of MATS.
  • Maybe there are strong near-term proxies for impact, like a panel of experts assessing a fellow's end-of-program research output.
  • Maybe we could just assemble a talent database of everyone working in AI safety and chart their journeys in a way that usefully indicates whether MATS is reliably counterfactual.

(Note that challenge trials is probably the wrong framing because the risk is not to MATS participants, but to broader civilization.)

Load more