If it transfers misalignment, we might get a misaligned model that’s easier to incriminate. If it doesn’t, we might get a capable benign replacement model.
Can you distill an aligned teacher model to create an accidentally misaligned student model? Ig I’m wondering how much evidence you get about the teacher model’s misalignment if the student is misaligned.
It seems pretty unlikely to me that distilling an aligned teacher model would produce a misaligned student model. (Especially conditional on us getting an aligned teacher in the first place.)
Can you distill an aligned teacher model to create an accidentally misaligned student model? Ig I’m wondering how much evidence you get about the teacher model’s misalignment if the student is misaligned.
It seems pretty unlikely to me that distilling an aligned teacher model would produce a misaligned student model. (Especially conditional on us getting an aligned teacher in the first place.)