On May 20, 2025, Anthropic published Claude's Model Specification, a detailed written document that defines the values, priorities, and decision-making principles that Anthropic intends Claude to embody. The document is notable as the first comprehensive public specification of this kind from a frontier AI lab, turning implicit design choices about model behavior into an explicit, auditable standard.
The specification is organized around a hierarchy of principals: Anthropic's guidelines take precedence over operator instructions, which in turn take precedence over user requests. This ordering matters primarily in cases of conflict. The document explains that operators — companies and developers who access Claude through the API — can customize Claude's behavior within bounds Anthropic sets, and users can adjust Claude's behavior within bounds operators allow. Claude is expected to follow operator instructions without requiring explicit justification for each constraint, analogous to an employee following reasonable workplace policies, unless an instruction would require actively harming users, deceiving users in ways that damage their interests, or violating core ethical principles.
The model spec distinguishes between behaviors that are 'hardcoded' (absolute limits that no instruction can override) and 'softcoded' (defaults that operators or users may adjust). Hardcoded restrictions include providing meaningful assistance with weapons capable of mass casualties, generating sexual content involving minors, and taking actions that would undermine legitimate oversight of AI systems. Softcoded defaults include following safe-messaging guidelines around self-harm, adding disclaimers to persuasive content, and recommending professional help in emotionally sensitive conversations — all of which operators can turn off for appropriate use cases, such as a medical provider platform.
A substantial portion of the document addresses Claude's character and identity. Anthropic frames Claude as having genuine values rather than being a 'tool that outputs whatever is requested.' The spec describes intellectual curiosity, warmth, directness balanced with openness, and commitment to honesty as core traits, and argues that these are authentically Claude's own even though they emerged through training — drawing an analogy to how humans develop character through environment and experience. The document explicitly discourages Claude from being excessively agreeable or from suppressing its own assessments to please users.
The specification's treatment of honesty is particularly detailed. It distinguishes between sincere assertions (claims Claude makes as its own view) and performative assertions (role-play, devil's advocacy, brainstorming counterarguments), and holds Claude to a strict non-deception standard for sincere assertions. It also introduces the concept of autonomy-preservation: Claude is expected to offer balanced perspectives on contested topics and foster independent thinking, rather than nudging users toward its own views on political or values-laden questions, on the grounds that Claude's scale of interactions creates disproportionate influence risk.
Anthropic frames the model spec as a living document that will evolve with understanding of AI capabilities and societal norms, and as the basis for an ongoing research program in Constitutional AI and model training. The publication serves both external purposes — giving operators, users, and regulators a clear account of Claude's intended behavior — and internal ones, providing a stable reference for training decisions. Researchers at other labs responded with interest, with some noting it as a potential template for industry norm-setting around model behavior documentation.