Constitutional Alignment: How We Keep Bloom Labs Models Safe
Building the most advanced language models on the planet means nothing if we can't trust what they say and do. From day one, Bloom Labs has treated alignment and safety research as a first-class engineering discipline — not a bolt-on afterthought, but a core part of how every model, from BabyMax 5.76 up through Homan 2.96, gets built. Today we're pulling back the curtain on the process: constitutional alignment, internal red-teaming, and the responsible scaling policy that governs every release that leaves our building.
Helpful, Honest, Harmless
Every Bloom Labs model is trained against a written constitution — a set of explicit principles the model is taught to reason about and apply to its own outputs, rather than a static list of banned phrases bolted on after the fact. The constitution boils down to three commitments, in priority order when they conflict:
- Helpful. Give the user a genuinely useful, high-quality answer, not a hedge or a non-answer.
- Honest. Never fabricate facts, never overstate confidence, and say "I don't know" when that's the truth.
- Harmless. Refuse to assist with anything that could cause real-world harm, and do so transparently rather than silently degrading quality.
Instead of hand-labeling millions of examples, we have the model critique and revise its own draft responses against the constitution during training, dramatically improving how consistently the resulting behavior holds up outside the training distribution — including on prompts the model has never seen anything like before.
Internal Red-Teaming
Before any model reaches an API key or a chat window, it spends weeks in front of our internal red team — a group whose entire job is to try to break it. That means adversarial prompting, jailbreak attempts borrowed from the wider research community, edge-case stress tests, and long adversarial conversations designed to walk a model somewhere it shouldn't go one small step at a time. Every failure gets logged, categorized, and fed back into the next training run. Homan 2.96 went through more red-team hours than any model we've shipped to date, and it shows in the results.
A Responsible Scaling Policy
Bigger, more capable models carry bigger responsibilities. That's why Bloom Labs operates under a responsible scaling policy: before we train or deploy a model meaningfully more capable than our last one, we run a formal risk assessment, define the safeguards required at that capability tier, and only proceed once those safeguards are in place and verified. Capability and safety are scaled together, on purpose — not one first and the other later.
None of this means our work is finished. Alignment is an open research problem, and we expect our techniques to keep evolving as our models get more capable. But we believe that shipping the most advanced AI in the world and shipping it responsibly are not in tension — they're the same job. For the full picture of our commitments, safeguards, and ongoing research, visit our Safety page.
This is, in plain terms, exactly why Homan 2.96 hasn't been released outside Bloom Labs yet. Internal development and testing are done, but a model this capable doesn't clear our responsible scaling policy on schedule or sentiment alone — it clears it when our safety team's risk review says it's actually safe to hand to the public. Until that review is complete, Homan 2.96 stays internal. We'd rather take the criticism for being slow than the alternative.
Read Our Full Safety Commitment »