Jonathan is exploring whether a model's implicit constitution can be inferred from its behaviour. He previously worked on AI control with Charlie Griffin through LASR Labs, and is pursuing PhD in chemistry at Imperial.
Jonathan is exploring whether a model's implicit constitution can be inferred from its behaviour. He previously worked on AI control with Charlie Griffin through LASR Labs, and is pursuing PhD in chemistry at Imperial.