·morgin.ai
Study finds “uncensored” AI models still avoid charged words due to safety filtering embedded during pretraining
Pretrain Forensics measured a behavior it calls the flinch: when a model avoids predicting politically or socially charged words even without issuing a refusal. Across seven pretra...
read →