Skip to main content
October 9, 20267:59

Can Haiku 5.5 Replace Decision Models?

By Samuel Gregory

About this video

Your AI budget is likely 5x higher than it needs to be because you are using the wrong models for simple decisions. In this video, I benchmark the brand new Clefft and Clefft Flash from Cloudflare against Jev and Haiku 5.5. We look at real world speed tests, cost analysis, and multimodal performance to see if the big models are finally being replaced. Key takeaways: - Clefft Flash is 22x faster than Haiku 5.5 in decision tasks. - Decision models like Jev and Clefft can cost 75% less to run than general models. - Haiku 5.5 remains superior for complex arithmetic and string logic. - Multimodal capabilities are now standard in smaller, open weight models. - The best architecture uses fast models as a primary layer with slower models as a fallback.

Your favourite AI model is burning your budget for no reason.

The world of AI decision models just got a massive shake up. While everyone was looking at the latest releases from the big players, Cloudflare dropped Clefft into the mix. This is a multimodal model that does not just process text but can actually run decisions against images. I decided to put it to the test against Jev and the new Haiku 5.5 to see if the premium price tags are still justified.

The Need for Speed

When we talk about decision models, latency is the only metric that truly matters for user experience. In my latest benchmarks, the results were staggering. Clefft Flash clocked in at a median of 90ms. To put that into perspective, Haiku 5.5 took 2.5s to complete the same task. That makes Clefft Flash roughly 22x faster than one of the most popular models on the market.

The Cost of Intelligence

It is not just about how fast the model responds; it is about what it does to your bank balance. Running 20 questions through the benching tool revealed a massive disparity in pricing:

  • Clefft Flash: $0.02
  • Clefft: $0.05
  • Jev: $0.03
  • Haiku 5.5: $0.23

Even when running Clefft through Cloudflare APIs rather than locally, it is significantly more affordable. However, the price of speed is often accuracy.

The Accuracy Tradeoff

While the smaller decision models are lightning fast, they start to fall over when things get complex. In tests involving arithmetic and string manipulation, such as counting specific letters in a sentence, Haiku 5.5 was the only model to maintain a perfect score. Clefft and Jev are fantastic for quick routing and simple logic, but they are not quite ready to replace the heavy hitters for high stakes reasoning.

The Verdict

Is Haiku 5.5 replacing decision making models? Categorically, no. But the strategy is shifting. The smartest move right now is to use Clefft or Jev as your primary layer and only hand off to Haiku 5.5 as a backup when your primary models disagree or show low confidence. This approach optimises for both speed and cost without sacrificing the integrity of your data.

Transcript▾

In my last video, we covered Tev and Nimble comparing them against Jev. But literally a day later, Clefft from Cloudflare burst out onto the scene with this multimodal model, being able to run decisions against images. And just yesterday, Haiku 5.5 designed for quick simple tasks that cost 75% less to run had me wanting more. That got me thinking, with these kind of improvements, does this now replace decisionmaking models? So by the end of this video, you'll know exactly that. And I benchmarked them all against each other. So I have the receipts to prove it. Geez Louise. Given KFT is open weights, we can see is they have a 27b model and a 9b model. Comparing those to Nimble and Tev, Tev being 4, Nimble being 9. It is a significantly bigger model. So I've built this benching tool which basically compares KFT, Cleft Flash, Jev, and Haiku against each other running against 20 questions. So, let's just see the results. I haven't run this yet. I don't know what the results will be. Let's go. So, Clefft steaming ahead there, answering the questions. Very well. Flash. Geez Louise. And Jev is probably probably the same as KFTF, but maybe a little bit slower. And then here we go. Haiku. Still fast, but clearly a lot slower than these decision models going. Here we go. Cleft Flash is 22x faster than Haiku. Here we've got some averages here. Clefft 266ms. Cle flash 90ms. Jev 282. So a little bit slower than Clev and of course Haiku being 2.5s as a medium. All of them getting the answers correctly which is great. Um and we can see the costing here. Cleff being at $0.05. $0.05 to run against those 20 questions. About half that for Cleft Flash Jev being even smaller. Really, really good pricing on it there. And then HiQ being nearly $0.23 uh cents for that entire run. It's worth saying that the prices against Clefft cuz we're running this locally which means it's completely free, but the prices against them are from the API prices over on Cloudflare's servers. And we can visualize that latency here. obviously cleft flash coming in very very fast there and then the costings uh can clearly be seen there. So overall I mean you know at the end of the day you're going to be running this against your own data sets the things that you want to be using but cleft flash is demonstrating itself to be a very very promising model given its cheaper price and a hell of a lot faster and whilst haiku is still pretty fast it just doesn't beat these decision models. So if that helped you make a decision on what decision model you should use then that deserves a like and subscribe. Now the biggest aspect of KFT or the most exciting aspect of KFT that is actually multimodal. So Jev the system 2 model might have multimodal or the next revision of Jev will have multimodal. However doesn't have it right now. So I've got an image benchmarking here that we can run against and then we can see the results of this. So, let's run this. I mean, Clefft being really quite fast on that given it's multimodal. Cle flash steaming ahead once again getting all of them right. You can see the ticks there and the highq model again just being that you know 2s plus. So, Cleft Flash is 6.2x faster than Haiku. Again, killing it. Under a second for KFTF, under half a second for Flash, and 2.2s for Haiku. All of them getting it correctly. And again, visualizing that latency and the cost against those. What I really want to see right now is a harder set of question to see it where these decision models start to fall over where it might make sense to go with haiku and justify its slower speed and higher cost. So let's just go through that now. So this time we're just going to run them all and see how we get on. So we're starting to see some false results there, which is good. Then we can start to weigh up where they're falling away. This is probably Jev then HighQ. trailing along. What I'm taking note of here is that the multimodal image is roughly I don't know half as slow if that makes sense. Like you know take half of the text response and add that. So there we go. Scroll up. Let's look at these. So order arithmet order arithmetic did the customer spend more than $100 in total during March 2026. Cle actually got that wrong whereas Flash got it right which is very very interesting. I think that's probably an anomaly there. We'd have to run these tests a few times to uh start to see what what's actually going on there. All of them getting this one where Haiku got it right. And I guess it's important to say that Haiku got pretty much all of them right except that one. So what is that one? How toxic is this message? So I guess that's kind of open to interpretation really, isn't it? Let's float those model names in the table so we can see who's who. Cool. So yeah, once again, how many times does the letter I appear in the text total? I would guess that's quite an easy one, but of course these are the um flash or decision models are getting incorrect whereas Haiku is getting it correct. Then looking at the images obviously again Haiku getting them all correct. It's kind of a mixed bag but um from a speed perspective generally we're seeing KFT flashes the fastest. It's the same story really. Jeff pulling out on average a bit faster this time over KFFT and Haiku coming in last but an overall higher percentage costing a little bit more. So a mixed bag really when it comes to all of these results. Generally again I think the bigger picture here is to run it against your data and have some uh some fallbacks in place so that if for things aren't uncertain or whatever that you hand it off to potentially haiku in a backup scenario if the decisionbased models are a bit uncertain or they don't you know they they disagree with each other. So there we go. I hope that answers the question. Is these smaller, faster models such as Haiku 5 Boom 5 replacing decision-making models? The answer is probably categorically no. But it's really nice to see the speed increases from them and make them or even more justifiable as backup solutions when the decision-based models disagree against each other. So, let me know what your decision model of choice is down in the comments. Like, subscribe if you haven't already. Until next time, keep on vibing.