About this video
Manual video editing is officially a waste of your life. In this video, I put the latest AI models to the test to see if they can actually replace a human editor for high quality podcast clips. We compare Fable 5.1, Qwen Max, and local models to find out which one reigns supreme. \n\nKey Takeaways:\n- Fable 5.1 is the undisputed champion of creative video editing and script comprehension.\n- Local models like Qwen 3.8 are currently too slow and unreliable for professional workflows.\n- Qwen Max is highly capable but tends to take shortcuts by reusing existing assets.\n- High intelligence in a model directly correlates with better, more logical video cuts.\n- AI allows for 'prompt injection' during filming to guide the future edit in real time.
Your Video Editor Is Obsolete And You Are Still Wasting Time
Manual video editing is a soul-crushing endeavour that consumes hours of creative energy better spent elsewhere. For years, we have accepted the tedious nature of cutting, syncing, and layering as a necessary evil. However, the rise of specialized AI models suggests that our time behind the timeline is finally coming to an end. This exploration into AI video editing agents reveals a startling truth: the gap between human intuition and machine execution is closing faster than anyone anticipated.
The Experiment: Local Versus Frontier
I put four heavyweight contenders to the test to see which could handle a complex, multi-cam podcast edit. We looked at Fable 5.1 (Claude), Qwen Max, Qwen 3.8 Flash, and GLM 5.3 Flash. The goal was simple: take two hours of raw footage and turn it into a succinct, engaging social clip.
The results were a mixed bag of technical brilliance and hilarious failure. While some models thrived under pressure, others proved that local hardware still has a long way to go before it can replace the cloud.
The Winner: Fable 5.1 Takes The Crown
Fable 5.1 emerged as the clear victor in this digital gladiatorial arena. It demonstrated a level of nuance that the other models simply could not match. It understood the script, identified the perfect moments to cut, and even handled the tonal consistency of the audio with remarkable grace. If you are looking for an edit that feels human, Fable is the current gold standard.
The Local Failure
On the other end of the spectrum, running Qwen 3.8 27B locally was a lesson in patience that I did not ask for. After forty eight hours of processing, the model had created plenty of assets but failed to actually produce a finished video. It is a stark reminder that while local AI offers privacy and zero cost, it often lacks the raw power required for heavy lifting in video production.
The Future Of Filming
This experiment raises a fascinating question about the future of content creation. Should we start filming for the AI? By injecting prompts directly into our speech or structuring our scripts for machine consumption, we can make the editing process even more seamless. We might lose some spontaneity, but the efficiency gains are impossible to ignore. Keep on vibing, and keep on automating.
Transcript▾
What started out as a curiosity on is AI actually any good at video editing turned into an exploration of different models local versus frontier and server side and how they compare against each other when editing videos. So for those that don't know I have a podcast with my friend Kabaza here and every single week we stream for two hours capturing the week's news in AI. We then turn those videos into these shorts with a mixture of screen shares. We've got full screen close-ups of each other here, all derived from one continuous 2-hour long screen. Now, this is very laborious. This takes up a lot of my time, and we've even chosen to really focus on the highlights of the week as opposed to just outright editing all of the videos and hoping one or two of them do well. So, I thought, why don't I get AI on the case? I took a look at the brand new Fable 5.1, Quen 3.8 Flash Max, the 27B version running completely locally, and GLM 5.3 Flash. So, let's dive into how they compared against each other, and we'll give my final verdict at the end. It all starts with this video use skill from browser use. I would say a lot of this is very opinionated from browser use or or video use. There's probably a lot of things that you might like or or prefer with your specific edits. However, I think this is a good starting point. Things like, if we go down here, sort of animation speeds, creative direction, typography, stuff like that, I think you could probably tweak to to your own. You also have a lot of these reference files that actually tell you about a bit of detail about those. Production quality is another interesting one. Now, to install this, you're simply just going to tell your agent to set up browser use. It will install it as a skill. However, you might want, as I did, I kind of want them projectbased rather than systemwide. So, I went ahead and I actually did it the manual way, cloning this repo. I did it the complicated way. And you will also need ffmpeg YouTube downloader. Could be for other pieces of footage or whatever. And it also tells you to use 11 Labs setting up an API key to do the transcription. Now, this is where I personally recommend Whisper. If you've got a newer computer, if you can run this sort of thing, it is a local AI, completely free. I would transcribe it using Whisper and pointing the prompt that initiates the video editing to specifically these transcribed files. So, I started out by running Whisper against the each of the video files that were downloaded from Streamyard. So, if we jump in here once again, we've got the screen share, we've got Cabaza's camera, and we've got my camera. And I've simply run each of them through Whisper to get these transcription files. You can see an SRT file, even a JSON file here, uh, txt files of everything that we're saying, including their timestamps here. And the prompt we're going with, obviously using the video use skill, the core video and screen share is this one here where it's both Kabaza and I with the screen share capabilities, the individual full screen, me and Kabaza, and then mentioning the transcript files that we created using Whisper before. Edit the web flow section. So I'm asking it to look at the script obviously. Take the web flow section and edit that. You have creative control. Create all edit and export files in a fable folder. Make the edit um succinct and engaging. Let's do that. Um and use design style from command.com. We do have a command design system skill as well, but this might not work inside of war, which we're going to we're going to use the other models. But let's put it in there anyway. Let's see how this fares. Okay, that's finished. But what I've also done is render out some versions using Quen Max, uh, GLM 5.3 Flash, and Quen 3.8 Flash. So, we'll go through all of these now and sort of pick out some of the the key points where I feel like they shined and where they kind of fell apart a little bit. And this one here is Quen 3.827B. Now, I'm going to get this right out the way for the start. This 27B was literally running probably about 48 hours and it still hadn't edited the video. I don't know if it was caught in a loop or anything like that, but you can see here if we go to videos here, it created a bunch of assets and things like that that it could use. Obviously, some Python scripts, which is all part of the video use skill, but it never ended up making a video. So, we can just rule that one out completely. This was a 4-bit quantization, so not the smartest implementation of Quen 3.8. So, I wouldn't necessarily rule this out. However, all things considered, I don't have an answer for how goodw 3.8 is at video editing. So, Fable here, take a look at this. If you need any more evidence that web flow, you know, we've got an intro here, which is quite nice. Software titles as well, not the way the other way around. And that says quite a lot. And a really nice intro slide there as well. Approaching the Webflow confics in less than two weeks. We workflow again is trying to tease us with what's coming up this time uh in 1987 and that was a nice edit as well. Looking at the overall time as well, 5 minute 57 it's quite a nice amount of time experience that they are talking about and I think it's going to be this kind like deterministic but yet very free UI that you can in there was a couple of edits there. I don't know if you noticed but it was it made sense and it flowed really nicely like the voice was even all the way through. You probably wouldn't even know if there was an edit if you weren't looking at the video. So, overall, really nice. And we've got these chapters as well. And it's cut to me speaking with title cards and all the rest of it at a really logical place. Web flow. Really nice ending as well. So, I think Fable's done an exceptional job there editing this video. If we go into Quen Max, both Quen models were quite hard to wrangle. Um, we'll see in this example here that looks suspiciously like the Fable edit. Quen tends to be quite lazy and that it's found other videos or other animations and just use those instead of doing what I told it, which was do everything from scratch. This is probably a unique use case for me because of course we've got all these experiments. it might find all these videos. If I was to run this genuinely sort of authentically, maybe it wouldn't have gone this route, but this is obviously what I found. And yeah, that laziness that Quen seems to have. Okay, intro is a bit sketchy there. Apple shipped something and gave it probably would have cut away to the article here. the same engine and I'm noticing that it isn't quite as tight as Fable with Hyperard you deterministic but yet very free UI that you can install or like add to your project based on whatever you need and this is what what they're uh trying to tell us and we will we will know on September 1st. So I gave kudos to web flow. Another nice edit there. Again, you're seeing these same tiles that Fable created. So, it's just stolen those, but the edit itself is different and seems to be pretty good again with a a reasonable amount of time. Now, if we take a look at Quen Flash, this is where I forced it to make its own graphics because it had stolen them as well. If you need any more evidence, the web flow the game could have been a bit tighter. No, no, can't do too be too harsh. Should mold and that wasn't a clean edit at all. Looks slightly different. So, I'm starting to believe this is it own graphic. Again, this looks like Fable something strange and gave it away free with every Macintosh. It was called Hyper. Not a nice edit there. Built by Bill, but it did try and edit it. It tried to edit that that joke that I made out. which I respect, but it didn't do a good clean job of it like Fable did. And drag and drop. We do have these sections as well. But you'll see here that this is actually taken a conversation from stacky because we mentioned web flow. I would have separated these out, but it's obviously not understood that it is a separate subject even though web flow was mentioned. So this is all pointless stuff from this point onwards. So in GLM here we've got one talk about it about the next web flow tease. They are teasing us as web related but probably could have been trimmed. I wanted to kind of like mix it with the comment section part of um the audio is out of sync completely creative things but it has created its own graphics which is good. H we how should we but the edit yes already seeing it really quite bad. I like that it's created these graphics though. These are interesting. These overlays. Yeah. I mean it's funny because he did Yeah. Awful edit, but the graphics are pretty cool. As for pricing, it's a bit of a weird one. I had to correct it a couple of times. So if this is my day here. So the Quen Max one cost 562. The Flash one cost 80. So, actually really good. The Fable one cost $15 if I was using API. Obviously, I have the Claude Max plan, so it doesn't matter to me, but that's just an idea. And overall, it took an hour, which is quite interesting. All of these generally took over an hour, which is uh surprising amount of time to be honest, just for one video. And you can see I've used 303,000 tokens today. This would have cost uh about 15, so pretty good price. Now, Fable 5.1 is the clear winner amongst them all. It seemed to understand a lot of the script a lot better than all of the other models. It seemed to pick up on what bits to remove. You're obviously dependent on the bits either side of the chunk that it removed that that from a from a tonal aspect, how how you're speaking is consistent to make it seem like a an edit that um that makes sense. However, given YouTube, there's a lot of jump edits anyway. You might be able to get away with it, but that's not really the the model's fault in that sense. I would be faced with the same issue if I was editing manually. The Quen Flash and the GLM Flash did an admirable job, but they don't pick up on a lot of the nuance that goes into a lot of the script and things that were being said. Quen Max was um a clear winner amongst the sort of open-source models. So, it's it is honestly about the more intelligence does make for a better edit. Although, that's my kind of um there's where my focus is on editing. It depends on the type of video that you produce at the end of the day. If you're focusing more on graphics or you rely more heavily on graphics and things like that, then you can of course prompt your AI agent to produce a lot more graphics. I think all of them are quite capable of doing that and following the direction from your prompt. I don't think I would be using this in my day-to-day. Although Fable 5.1 is something I would be interested in trying to shape to make it a better editor. It does raise the question on actually filming in a specific way. The production in of itself is the is recorded entirely differently. Maybe you have a script. Maybe it works better on scripted videos where you are saying everything that needs to be said. Maybe you you say notes actually in the script. So Claude put a graphic here of an XY Z and then go back to the script. So there's actually prompt injection happening in the script knowing that AI is going to be editing this. So there's an argument to be said about changing the way that you record these videos knowing that an AI is going to actually edit them. I think you just lose a lot of fluidity. You lose a lot of uh creativity and you lose a lot of those human elements. I mean this is not exactly a surprise. However, again, all three performed technically perfectly. I guess uh it's that creative aspect that they differed and Fable I would say arguably was the most uh accurate and the most creative out of all of them, but it's also a lot more expensive. So, equally, Flash is a lot more is is very impressive in its own right. Obviously, I use AI to code. I just wanted more interesting ways or different ways to test how different models work. This was a really interesting use case that I haven't yet explored. If you're using your AI for a really interesting use case that you'd like to know or you'd like me to test certainly from if we can get local AI involved in all of this, then we should absolutely try that. Let me know down in the comments. Like, subscribe if you haven't already. Till next time, keep on vibing.