Question: is it good to work on better consumer/enterprise AI applications so that the public can get a better sense of how capable AI systems are?
Paul Christiano on the impact of RLHF research.
- He thinks it was good. Downside is that it increases investment in AI capabilities. Upside is that these capabilities already exist in the model, they're just latent, discovering them later would lead to really drastic accelerated investment.
- The hardware overhang argument doesn’t work: better products means more investment means more hardware means more overhang…
RLHF (and broadly making models more controllable) made LLMs much more useful and led to a lot more investment for AI. I think that's bad, given we don't know how to control AI systems. There's now a lot more incentive to make them capable and obedient.
I think building really successful AI applications is also bad, though to a lesser extent than core research. RLHF increased investment by drastically increasing the amount of inference people bought. Really good applications also increase the amount of inference people buy, which means more money for capabilities work.
One argument for building good applications is that it helps people realize how capable AI systems are, and this will help people better understand the risks. I'm not sure about this; companies and governments won't want to slow down a money-printing machine. I think the better approach is to just avoid increasing more investment into AI. More money means more capabilities and practical alignment work to make AI systems seem safer.
What about narrower AI systems like AI for drug design or a really good virtual cell?
- Capabilities research tends to be transferable across domains: AlphaFold 2 used a transformer and AlphaFold 3 used diffusion. One can imagine a world where protein ML had discovered these techniques and then they transferred to language to create useful general AI systems. So capabilities work in other domains is accelerationist.
- Some real examples of domain-specific → general transfer: residual connections and convolutions came from vision, attention from machine translation, RL from games.
- But again the link is weaker than RLHF. RLHF motivates direct investment; ML for bio might cause more people to focus on and invest in AI which means more people do more transferrable research which speeds up general systems.