I like MATS (Machine Learning Alignment & Theory Scholars). It has produced excellent research and made valuable contributions to AI safety.
Still, I see a worrying trend: some MATS research appears designed to portray Chinese open-weight models as "evil."
Anthropic helps shape part of MATS's research agenda. Anthropic also helps define the evaluation standards through its researchers, concepts, and model-based judges. This influence deserves scrutiny, especially when Anthropic is a closed American AI company competing with Chinese open models.
Terms such as "evil persona" are anthropomorphic and normatively loaded. Anthropic's persona-vector research first defines "evil," generates opposing examples, and then extracts a direction corresponding to that definition. The result inevitably reflects assumptions embedded in the experimental design.
There is also a structural asymmetry: Chinese open models are altered and dissected, while closed American models rarely face equivalent scrutiny. American models then often judge the resulting behavior. Technically valid findings can still receive geopolitically biased interpretations.
Stronger research should apply identical tests across model families, use independent human annotation, include culturally diverse evaluators and multiple judge models, and clearly distinguish an original model from a deliberately corrupted derivative.
MATS does excellent work. Greater methodological symmetry and cultural neutrality would make its research more credible.