Anthropic concluded in the April Mythos Preview alignment risk update that the model “does not possess any unknown propensities that would increase alignment risk.” The report argues that if Mythos Preview were coherently misaligned[1][2], it likely would have been detected by the assessment (following Anthropic, I will call this “reliability of the assessment”
> Current assessments aren’t critical because Mythos-level models seem unlikely to cause catastrophe to deploy on priors.
Is this a mistyped sentence or am I misunderstanding
I don't believe it's mistyped. What do you think we disagree about here?
I think it’s a grammatical error. “To deploy” is dangling in this sentence and I think you didn’t mean to write it there
Claude explains in more detail
https://claude.ai/share/dfc1b318-e212-4785-ab01-f3873d40e792