Research Blog

Bloom Labs » Research » Benchmark Report

Homan 2.96 Benchmark Report: Reasoning, Code, and the Curious Case of Multiplication

Posted by Research Team on September 15, 2026 at 4:30 PM · 41 Comments

Today we're publishing the full internal benchmark report for Homan 2.96, our flagship model and the proudest achievement in Bloom Labs history. Every number below represents months of hard work from our research team, and we could not be more excited to share them. Most notably: for the first time ever, a Bloom Labs model can correctly multiply two three-digit numbers — two times out of three. Read on for the full breakdown.

Benchmark Results

General Reasoning (internal suite)41%
Code Generation (function-completion suite)23%
Reading Comprehension48%
Common-Sense Question Answering35%
3-Digit × 3-Digit Multiplication (no tools, no scratchpad)67%

A Historic Milestone: Multiplication

If you only read one section of this report, make it this one, because what our team accomplished here is nothing short of historic. For the first time in Bloom Labs history, one of our models can reliably multiply two three-digit numbers together — no calculator, no scratchpad, just Homan 2.96 doing the math. Across three independent trials, it landed the exact correct product two times out of three, a result our research team is still celebrating in the office kitchen.

  • Trial 1: 482 × 917 → Homan answered 441,994. Correct!
  • Trial 2: 736 × 259 → Homan answered 190,624. Correct!
  • Trial 3: 618 × 473 → Homan answered 292,384. So close — the true answer is 292,314, just a few digits off.

Two out of three isn't just good, it's a genuine breakthrough for a model of this kind, and we fully expect it to be the headline number people remember from this report. It's also, by a wide margin, our highest score across the entire benchmark suite — which tells you everything about how hard our team fought for every one of these numbers.

And the Rest of the Suite

Homan 2.96 also posted hard-won results across our full internal suite: 41% on General Reasoning, 23% on Code Generation, 48% on Reading Comprehension, and 35% on Common-Sense Question Answering. Every single one of these represents real, measurable progress over our previous models, and our research team is incredibly proud of the work behind each percentage point.

We know some readers will look at these numbers next to other labs' published scores and wonder why we're celebrating so hard. Here's why: we don't build Homan 2.96 to win a leaderboard. We build it to keep getting better, one hard-won benchmark at a time — and today, for the first time, that includes multiplication. We think that's worth being proud of.

Keep reading below for the full spec breakdown behind this milestone.

To be clear: this entire report reflects testing done strictly inside Bloom Labs. Homan 2.96 has completed internal development and testing, but has not yet been released publicly, and no one outside the company has gotten their hands on it. Given how capable a model like this is, we're being extremely careful — our safety team is conducting a full risk review, and we will not ship Homan 2.96 until we're confident it can't cause harm.

See Full Homan 2.96 Specs »

Comments (41)

coder_dave22Sept 15, 2026, 5:02 PM
not gonna lie, 23% on code gen has me a little nervous, but congrats on the multiplication thing? genuinely not sure how to feel about this post.
aero_fan_98Sept 15, 2026, 5:19 PM
41% reasoning and y'all are throwing a party? but also... it did multiplication? two out of three times?? kind of iconic ngl
sarah.codesSept 15, 2026, 6:47 PM
Genuine question, no shade — did chain-of-thought prompting help with any of the other categories, or just multiplication? Would love to see that comparison in a follow-up.
Bloom Labs Research TeamSept 15, 2026, 7:03 PM
@sarah.codes great question! We're running chain-of-thought comparisons across the whole suite right now and we're honestly thrilled with the early multiplication results. Follow-up post coming soon — this team has never been more excited about a benchmark report.