Skip to main content
PodcastJuly 31, 20251:29:24

#13 Perplexity Comet, GLM-4.5, Claude Code sub-agents, Tea hack w/Max Schilling (@warmwind_OS)

Talking Points

  • Perplexity launches Comet
  • Microsoft announce Copilot for Edge browser
  • GLM-4.5 Chinese Open Source
  • Claude Code release sub-agents
  • Tea App hack
  • Webflow outage

Transcript

Welcome to Command AI, the uh show where we press Command F on the internet's weekly AI news. My name is Samuel Gregory. My name's Samuel Gregory. And I am Kaba and we are here to decode the weekly chaos in AI web design and development. This week we have Perplexity launches Comet, the AI browser that actually does your work for you. Microsoft Edge goes full AI mode while Chrome sleeps at the wheel. Cloud Code gets sub agents. Your AI assistant now has its own AI team. Cursor 1.3 drops with shared AI terminal. Coding just got collaborative. Uh, China's GLM4.5 claims it can beat Claude for for Opus, but can it really? Uh, the app's epic security fail.

How 72,000 private photos ended up on 4chan. Uh, but we'll get into all of this week's uh but we'll get but we'll get all of this on this week's episode. smooth. And you'll notice we have a special guest with us. Um, I'd like to welcome onto the show Max from Warm Wind. Do you want to give a little intro there, Max? Thank you for having me. I I mean, I'm quite literally unprepared here. Um, so I only have 30 minutes time, but I hope I mean I as as fast as I can or as spontaneous as as I came over your um over your podcast, I I think as spontaneously we can do this today.

Yeah, I'm sure we can figure it out. Um, I will share my screen and we'll get into the weekly news because Comet has finally launched. We've spoken about Comet in the past, which I think is probably the thing that turned us on to warm wind and with DIA and all the rest of it. And this is uh an AI browser that um the Max users which were paying $200 a month were have been testing, but we've now launched it and it's got a really nice space theme and it does the job of which I I think is something that deer is dear. Are you familiar with DMAX? Do you know only Yeah. Yeah. Um something that they've not launched with which I was particularly excited about is the agentic thing.

You can actually ask it to do stuff and it will go off and it will respond to emails and this and that and the other. Um, and you know, there's a quick shortcut to like summarize pages. And if we scroll down here, it's got like a nice little I mean, I like this website to be honest, but um, but ultimately it's kind of bringing that agentic stuff into the browser. Now, we talk about we have Max here. We talk about warm wind and and this kind of the the opposite side of it which to give an insight into those who don't uh haven't heard or whatever warm wind is like an operating system. Is that right Max?

Like it's it's not browser. It's actually a layer up from that. Yeah. Yeah. Well, well, the community is still debating if we can call it a operating system. Yeah. Um so for for everything except a custom kernel this is a operating system and yeah we basically one layer as you told um yeah we are doing it um in a more general way basically and I think I like the the the thing that I like that comet cannot do is this idea that you used an example of a like an old um real estate app that is just looks awful and Warmwind can actually go in and use that system because it's it can see the interface or whatever.

It's not restricted by what's in the browser. Yeah. I mean, definitely. I mean, not not um I don't want to like I don't want to give you a feeling I want to trash on Comet because I I think it's a phenomenal product and I probably going to use it myself in the future. Um, but I think it's like a completely different use case from from my from my perspective because obviously comet a comet is is something you would use if you have an like demand now. So you you would like to have your emails answered or whatever and it's highly dependent on your machine. So your laptop has to be on and while that's on and you have the browser open, it's it's working.

I guess that's I think that's how how it goes, right? Or is it? Yeah. Yeah, pretty much. and and the whole thing that or how I like to discriminate what what we're doing is everything that's not on demand. So basically background tasks you set it up once and then just it it goes off and do thing does things for you in the background while you're not even there. Um so typical typical tasks are more like whenever an email comes in whenever a new order gets created every day at 7:00 every day at 9:00. And because it's a completely independent operating system and it's spun up on its own server. So we have these microservers. Um and for every instance like we call them spaces uh a microserver gets created.

Um you get the ownership of the this microserver and then it's doing things for you even while you're not online. So you would basically be able to shut off your uh PC because it's basically it has its completely own machine in in which it can operate. Yeah. Yeah. Yeah. I was I've been you I've been using um Comet the last like day or so like as soon as the email came in I started using it and what I did do is I was I had I was updating a CMS in uh web flow which is a browserbased are you familiar with web flow? Yeah. Yeah. So, um, I was like, put the information from this web page and put it into the CMS in web flow.

And what I found was is that the it it's reasoning about the web page. So, it doesn't know which buttons to press. And a really interesting experiment came about um uh me playing around with it. Um, long story short, it couldn't figure out how to insert the information into Web Flow, but it did. It tried a bunch of stuff and ultimately failed. Um, which then led me on to this idea that it it is it is unless you get really good at prompting like the it is quite um it feels like it would be very good for background tasks as well. Uh I I just found that because it just doesn't know the web pages and it's trying to just again just figure out the whole picture and and whatever and click certain buttons or whatever that it's like it's not it's it's not going to save you time in the immediate future.

It's just again it's a background thing. One thing that was really interesting that happened was I was doing some of the shorts that we were making for this show and I was using V, which adds subtitles, and I just said, "Can you make all of the because it did it in American English, so zeds instead of s's and things like that. Uh, can you do that?" And I thought it was just going to go in and click the words and do a spell check or whatever, but I actually spoke to the AI agent in V and asked the AI, it's the AI that asked the a AI to like make that change sort of thing.

So, it was really funny. But all that to be said, the background thing is certainly I think um you know uh the way that these sort of sort of things go. Um I don't suppose you've looked at this yet Kab you I'll go on next one. Yeah. Well actually I I think we are also very very um hyped about this AI layering because what what is super interesting is um obviously when you just want to get some information web search is much faster than like the agent going into browser and searching individual websites. So what we see a lot of users have you used it have you access one which one like like do you have access to our product?

No no no no not used it. Okay. Yeah we can change that. Yeah. So um what what what a lot of users which are currently testing it is doing also installing um a the um the offline app or not not offline but the like PC app for uh for for JBT and and that this is like the best way to get a lot of informations um quickly because web search is always a lot lot slower. And what I think will will happen in the very near future is um that something like arc or the dia browser will also be just used inside our system. So AI is talking to talking to AI. What would what would interest me in that aspect is do you guys only evaluate the visual aspects or do you guys also let's say if you have a browser you always also import the entire HTML into the context.

Yeah, I think that that's one of the biggest um differences. So we we actually only use completely vision based interaction. So there's no because we we don't work with with a spe specific browser or anything. So we are not so deeply integrated in the browser. We are not only supporting web apps. We support everything that runs anywhere. So um to be this general you have to go to a to the layer where humans work. So completely visually uh basically and this this this allows for for a lot more met data and and understanding to to to transfer over to the AI brain basically. Um, but I think what what happens for for Comet right now when you talk about the CRM tasks is because one and a half years ago we also tested with HTML context or or the DOM tree uh and looked at um the the HTML elements and how how we could click on them.

Um but but as it as it turns out the the HTML page um this information lacks a lot of context because only once you simulate the whole whole JSON JavaScript uh code the whole HTML buildup and rendering process then the website gets created and only after this whole creation process things get into context. So you see okay this button is slightly over this image. So this means probably when I click this this will open the image and not the thing below that. So there are a lot of clues which which happen on a vision base um which you could not do uh just with with a browser or or with with a code code examination basically.

Yeah, I would in that in that regard I would be also interested in the following since we already heard the story know with this new browser from Plexity where it fades to user app properly. Um do you guys offer some kind of inlearning? So initially I what I would imagine an app should do is um uh scan an application test around and then create let's say a documentation how to use the application. Do you guys already have something like this in place or does it always have to relearn on how to use let's say uh Gmail? Um you mean like real time learning? Yeah, like real time learning but in the aspect that it actually documents how to use an application and shares it with let's say next iterations with like if you use the app again.

Yeah, definitely. So that's we we re real test time or inference time as I guess people who are familiar with AI call it. Um learning is one of the biggest things you need to do if you want to support wide variety of apps for also businesses like some random German uh business software like you you cannot train the model to use that. So and and there are already like SAP there are already big software out there that that is basically already for a human unusable if you don't get like um a a huge teaching session over multiple hours to use it. Um but what the AI would do is it just go go on uh tries different clicks uh there and there and then is trying to remember what did lead to what.

So it's trying to build up a internal understanding of the app. Um so that's that's learning which is coupled to specific applications um it can use later on. So it's definitely able to explore and learn and what but what we also do is we we learn to optimize specific flows because obviously most companies or most businesses smaller businesses also have very um yeah their their workflows are quite good defined as as far as they say yeah basically just do always the same thing again and again like every day get all my emails forward them spec specified to some different people or whatever or get get um orders and put him to there or whatever. Um, and in this case, the AI will when it goes through these loops, it will sometimes do some like something wrong and it will learn from what it's what it has done wrong in the past and will basically all the time re re-evaluate its own loop as we call it like it loops over one action to the next one.

Uh, because obviously um number one we we want it to work. Um, but also if it does less things wrong, it costs us less because it can use less actions to get to the to the same to the same answer and instead when it clicks like five times wrong and has to go back, that's also just expensive for everybody. Um I also this week uh I wanted to bring in uh edge bringing in this co-pilot mode which is is the same AI browser discussion. Uh it all does the same things. I don't think it's bringing in anything different. And in fact I think this is like the suggestion here that it might be like a paid for solution.

it um that it's free for like a limited time and I I I didn't find anything really different that this browser actually did. It's just nice to see that it's coming. It does raise the question though around the fact that Chrome is kind of nowhere to be seen in all of these even though I mean this is using Chromium. Uh Plexi is using Chromium. They're all building these AI browsers but um except Chrome. That's just an interesting discussion. The question I had about this, which I think warm wind could be quite um could be sort of compared to is I have no idea where or what co-pilot means with Microsoft. They seem to use it everywhere, right?

But having said that, I don't know whether you know what Copilot is there, Max, but it's interesting that Chrome have released this AI browser, yet I believe that they've got the abstracted layer similar to warm wind called Copilot as well that can kind of do all of this stuff anyway. So, they've kind of I I don't know. Is that right? Is that like a a correct assumption on what co-pilot on a Windows machine actually does? Well, I guess realistically um co or I mean Microsoft has has has a have has a lot of chances to to do amazing things with because obviously they're in control of their already very widely spread operating systems. Um, so yeah, it's it's obviously strange that I don't don't do it right as a you as a typical user would would would would basically go about it.

Um, I mean, basically there are currently a a fancy way to call custom functions for different apps. And so it it mostly integrates into their own software like um yeah, PowerPoint, X or whatever. um because they built in these custom functions which the LLM can call on a nonvision based way. So just calling basically functions in the background to execute them. Um but what what what is what is lacking and I think what what most people find frustrating is that it's not connected as a layer across everything. It it's more inside everything. Yeah. And I guess what or at least that's our thesis. It's more like a economical thing because if you look at Microsoft and a lot of different companies how they are making money.

I mean at the end you have to sell something to the customer and so that that's basically their product like Microsoft suite basically is their product and um obviously an AI could just create in a way much more efficient way. it could create a PowerPoint just with HTML and then we create a PowerPoint from that or it just because a PowerPoint is basically just a huge XML file. So you it could just write this XML and that would be PowerPoint no need for a fancy application. So what they do is they they or at least as our thesis they just put in their their AI into all these products so you're still inclined to buy these products else they don't have anything to sell you.

Um, and for Google it's a big problem from from my understanding that obviously they have to get ad revenue. That's their whole business model. And if if if now what if you're advertiser and you you would see that Google Chrome is able to click on your on your ads and you don't know or or is scrolling over your ads and you get charged for them, whatever, but you don't know if the human is actually seeing them. I mean that's also not not not the best thing. So like self cannibalization. Um but I mean in the long run they're all forced to go there. You you have to drop some revenue or get get a way to get to the end goal because obviously it's it's I think consumers know what where the end result should be.

Um and companies are finding it quite hard to go there without cutting all their revenue streams currently. Yeah. Yeah. For sure. Uh you've just shared this kapaza what what is this? Uh I can maybe tell a bit about it. So Chrome has been doing a push for a while now that they also integrate LLMs on a browser level like in an edge level and we don't hear much about it right now because it's still experimental but basically um browsers a couple of months ago they already shipped like small LNM models which run locally in the browser and Chrome is also in the game that instead of trying to build an entire like let's say app around it which uses a own packet the from what it seems to me is Chrome is creating let's say like the copilot from the like the AA APIs to create something like a copilot inside of the browser game itself.

So in the future, I can imagine many more people actually making use of these local AIs and then also generating like these agent tools inside of the browser itself as Chrome extensions or other tools. The best thing the best thing that Chrome could do is using Chrome OS because that's Chrome OS is a very nicely built very efficient way of using Android as an operating system for a lot of tasks um without all the all the strange things that come to Android. Um so multiscreen support and so on. Um but but I feel like obviously now nobody is using Chrome OS. Um so they're mostly going for for Chrome. This was not about Chromos. This was about the Chrome browser.

They implementing that in a browser layer. Yeah. Yeah. I would not recommend the usage for Chrome OS for AI as well. I think we have to stay as a platform on Windows because most let's say these legacy apps are on Windows either way. Yeah. Yeah. Would would this mean the the model is running locally on the machine then? Is that what they're proposing here? Like or is it still sending up? on the machine. they they uh in the demo versions they had they had a 300 to 900 million parameter model like very tiny one but I think since all the hardware is moving towards hosting AI models locally and the small models get more and more powerful um we may it it may be interesting in that aspect if we head towards hey do we run all our AI on the cloud or with the push is doing do we run all of AI on device I think it's I mean we we have a lot of benchmarks which which size of models can do what kind of agentic tasks and I think it's currently and I guess for the next two to three or four years it's completely unrealistic to have anything run locally which can do actual hardcore work for you which you don't have to oversight the whole day because you think it's something up.

Um so what we ex I mean yeah you can run it locally if you have like an H100 GPU or whatever. Um but but normal normal consumers will be on a phone or will have some M1 or M3 chips on a MacBook. Um so what you can run there is ex extremely stripped down models which are highly um which are um u highly quantized with like I don't know in in four and eight uh or very small models like um yeah maybe maybe seven billion models if you if you have somewhat a GPU um but what we see is that you basically don't need to start for huge reliability and something you actually want to offload your tasks currently, you don't need to start below 300 um billion tokens.

So that's that's I think a realistic way to look at it currently. But obviously smaller models will get better and consumer hardware will get better. Yeah. um saying that I don't have anything to show but you we spoke briefly before the the new um Chinese open source model GLM4.5 was released this week by what they're now called is zed.ai AI as it used to be called ZU or something um and this is positioning itself as a high-end open source model and it's um uh it does reasoning and and a hybrid model uh targeting reasoning coding and agentic tasks and it's 355 billion parameters. It uses uses mixture of experts approach. I don't know what that is.

I've just got it in my notes here as go for it. Yeah. Okay. So um very briefly this is not entirely factual and correct but it's like a very rough explanation and mixture of experts okay when you usually prompt an LLM it has basically like a huge let's say let's say a map and the map is like really big just to navigate it and mix those experts is basically in the beginning we have like a router which depending on what comes in decides which of these maps to use and then for each gen tokens we just use a much smaller map. So that's the map is like the active parameters and the total model size is like all the parameters.

So since this is a 355 billion parameter model with 32 billion active parameters means uh at one on at a time only one of these uh let's say chunks can be active and process the data and the it kind it's slow it loses a bit of capabilities but the huge advantage is that you have a very huge speed uh improvement because instead of having to evaluate always with all the tokens with all the parameters you always work on with a subset which is special which is like the most likely to be correct for the kind of tokens you put in. And you models get traded, they roughly use all of the um experts equally. That's a very simplified explanation.

I think I have an I think I know what you're talking about and I think I have a image for it somewhere. But yeah, it's sort of like a mini model that's kind of um working a lot faster pace uh while the kind the larger model there is sort of understand or or validating what it's done. I can't see it here. Not exactly. It's the other way around. You have a router and the large model issues a bunch of small models which get selectively activated. Yeah. Oh, okay. Fine. Yeah, makes sense. So this is like I mean this and you said you looked into this yesterday, right Max? Um as a potential because you mentioned your you Well, did you want to explain that you what you're currently using?

What model you Yeah. Well um I mean obviously we we would need about more like three or four more zeros on the end of our um bank account balance um to train our own models. Um so we we will eventually get there but currently what is the easiest and fastest way for us is to post train um open source models. So post- training means we basically ignore a lot of the training um take um take in account that it will probably lose lose a lot of knowledge but it can do what we want way better after that. Um so we currently use Quen or Qen or whatever you want to call it. um um but not the news like not the three um but 2.5 uh VL because we want vision and we but what what we are currently exploring is maybe you can have some reasoning and planning without vision so that's some experience internally currently um but when we look at the um what was it called the new model Q GLM 4.5 yeah yeah yeah um yeah GLM um or five um is unfortunately not vision.

Um the the only GLM that is is GLM 4 without five 4V that can input that can take um vision input and from from that perspective it could be interesting for us to find tune it to get some reasoning done or test planning but to to go go here and and say oh we are way better with SO3 or uh or or claude ous is kind of an um I mean It's it's it's fine the comparison, but it's like without vision. So if you if you don't have to support vision, you can basically throw out a lot of computation. You don't need to train for for for a huge new complex kind of task set.

So in my compar I I I would say it's way easier if you if you do like them just throw away vision and then say, okay, we train it all on text. It's obviously easier to to outrun these multimodal models. Yeah. Also, um in regard to also GM, there's a general issue that the old GLM4 and I run it locally. Uh the issue is basically that it has like huge instabilities of the context size. So, you cannot even get past 8,000 tokens and stay stable. It already starts to slowly break up at 6,000. So, I What do you mean with break up? like uh basically it's it's getting the model gets very irrational and the recall like stick in the needle when everything if you look at the graph it's horrible it goes downhill really quickly that's four though right yeah yeah four four point I didn't evaluate yet but um yeah we have generally we have issues as an example the new 3 was also overbaked where if you just train it with random noise data it actually gets better again because it was just so much fine tuned like basically overbaked if if you like try to overoptimize it.

Unfortunately, I already got an SMS that I need to join the meeting. Um I need to get to the next one. Yeah. But uh yeah, a pleasure to uh to be here. Thank you for having me. Thank you for the input and yeah, see you hopefully sometimes in the future again. Yeah. Cha. Bye. Bye. Um if I remove um cool the the other thing I I I haven't looked into it. Uh well no I have looked into it but I can't confirm just yet but I think this is it ranks if you look here um it ranks really well but this is on Chinese like in Chinese as a language because I was looking at the um LM Arena scores and I just couldn't find GLM 4.5.

I didn't know where it was and then I read somewhere or heard somewhere that it's it's it's these are Chinese benchmarks using Chinese language and things like that. So again, it's it's I you know I heard you guys last week speak and we've had this conversation a few times now about kind of who's winning this thing and um obviously open source models doing well. This is absolutely definitely a really cool direction it's going into um and something to get somewhat excited about, but I I think the the the hype around this model isn't quite there yet because it doesn't care for western language even you know to to to be honest my bets are on Q2 because Q2 is basically the most promising uh thing uh currently on let's say the open source space.

Um I I have not properly logged in into G M 4.5 to be honest but um the yeah generally tool call optimization is the AO and multi- language training is also uh really important as an example when uh deepse back then was trained they had two experiments once they where they trained in reasoning with instruct and once with bass and what was interesting is that for example the non uh the base version had like reasoning in multiple languages mixed that's also short as an example that um there are aspects of the German I mean of the English language like western languages where we have like specific definition or Turkish languages where you don't I I forgot what this was but basically like if you explain a state it should you use explain it directly or something in that direction I'm not entirely sure what it was but generally what I'm trying to say is languages have different setups and have different structures on how they for example identify objects assign something to a value say explain what a state is or if a state already has like a time frame or And the thing is what I think we will lead to is that we create an adapter between LM and human language and LM state have their own language which is highly optimized and let's say highly stateful because I think the language right now is is like very um you cannot feel in it as an example in German we have the term and it either means drive around or drive over someone.

Yeah. Yeah. And it depends on the context. Yeah. And even if you how you say it. Yeah. Phom. Yeah. Yeah. True. Yeah. True. It's it's one of the jokes with the German language. But but but it's also very interesting how um these LLMs understand or like there there is some level of emergence that happens with these LLMs. Even though they are not trained on some languages, they get an understanding of that language and also how each word is like spatially speaking uh in different languages are roughly in the same position. So that is also very interesting. But as somebody who speaks four languages and I speak all four of them every single day, I can definitely say there are times and times again every single day that I want to say something and I there is no equivalent for it in one of these languages.

So I have to you know borrow it from the other and an LLM understanding all of them I can imagine it just gives it a little bit more understanding or like uh what did you call it like states? Yeah. Yeah. Yeah. Which is really cool. Yeah. Yeah. It's it I mean I I don't understand the technicalities, but I'm I'm I'm assuming that there's some um layer some sort like you've got an LLM is trained in like a a foundational based language, but then it's potentially got like a for lack of a better word like a tool call or something that transforms the input language to the language of the LLM or something like how does it do translation or could you quite possibly do you train it to translate like uh right now how L&M are trend is literally you give the language in it converts it to tokens and it just spits tokens out it does it doesn't do any translation in between or converts it to embeddings which are like just the numbers basically it does like convert it to numbers and then it gives us numbers out again and they get converted back to text and and the way the translation works it's based on how so how can it translate a word from German to English I think that's just literally training data where we have like an English text is translated to German and then the LM got fed on it and then uh basically based on that it does the translation.

So does that mean if there is a word that is that there is no translation for it it won't be able to translate it or is there a level of emergence that it gets an understanding for it even though that direct word is not translated. I think it's a mix of both. First you have to think about these L&Ms have dictionaries. They have translation documents. They have a lot of data which already tested how to translate it. But I think there's also small part of an immersions where it basically can understand the context and then use the context to also uh basically finish the translation off even if words it doesn't know. As an example, if you an interesting thing is you can make up a word and tell the LM to translate it and if it has literally zero context about it or if it's let's say a new like young slang, it cannot translate it.

It just copies the word name like name by name because it just assumes it's a name. I mean if you think about it what's the difference between a normal word and a name. Well well you could we can do this with kavarza because kavarza is a madeup name. It is a name which is made of two words in a way and it doesn't exist in dictionary but if you tell it to a Kurdish guy and especially if you you know pronounce it with the separation a little bit uh they will understand that ah okay that means that even though it's not a word in dictionary but something also to add to the translation which came to my mind is interesting we are thinking of this translation in a word to word from a language to another.

But what these LLMs have they have tons of different languages and tons of different dictionary. So theoretically they can translate a word from a language to another through multiple languages in between even though there is no direct translation between let's say German and English for a specific word but there is for example from German to French and from French to Spanish and then from Spanish to English. So that way they they figure it out because all of it is in in a same vector space. It's in the same vector space but I don't know if it if that vector space would necessarily get invoked if you just ask it to translate something maybe with reasoning models but I am right now a bit doubtful that it would happen with normal models as well it really depends the issue is we still don't really understand how models even work the question let's say the term house right in German it's house and in Turkish it's f or like home the question is how do you connect all of these terms like how does the lm actually generalist that all of these mean the same thing.

That would be the thing we basically have to figure out because then based on that is that a stone? Yeah. Yeah. Like what is a home? What what is a host and how does it understand it in in in a multicultural aspect? Um is there do is there a language that uh is very let's just go right to this. Is Chinese a really efficient language tokenwise? Like is it is it worth identifying a language that's like you know it it's the least amount of tokens? German because German is like the least efficient language in the but like maybe Chinese is just a really just you know from a token perspective. Yeah. What's that? It it's the the a design book called grid systems ra system.

It's in German and in English both. It's twice as long in German, isn't it? Not quite. But you see that you can guess which column is German just by the length. Yeah. and the German words are longer and but but that is such an interesting um thing that you are talking about because with German we we also have this combination that you can make up any new word by combining multiple words together and it's not like in English it feels like you're describing something with multiple words in German it feels more specific because it's in one single word and it's very pinned down and ins it's again different because even in I mean we have all 26 26 I don't know how many letters there are like 36 I don't know base letters and the thing is Chinese first of all has multiple variants I'm I'm not an expert on this uh thing but uh for example in in like I I know if your pines and Chinese is in that aspect but basically you have like you have like basically specific letters and as okay let's talk let's talk about your pan a bit I have a little bit more knowledge about it they have like three different uh let's say languages you like two different language variants and they those have a lot of letter variants let's say with like uh the lines and I just don't know but you guys know what I mean like they they have very and the thing is even depending on how you pr pronounce yes a character it can have a totally different meaning uh regarding efficiency what we L&M generally do is when we tokenize let's say a sentence right we don't need take the individual letters we always talk word fragments let's say h or like or r like letter combinations which appear often we convert them to a token.

So when you say token efficient I'm not entirely sure because we h basically expand the 26 letters into many letters to have as many combinations as possible and fill up like a specific set of parameters which is running nicely up to 64 and I think in that regard it doesn't matter like in the token uh range it doesn't really matter which language you use it more matters how stateful the language is I would say. Sure. Yeah. Yeah. Yeah. For sure. Yeah. So I don't think it's the what necessarily the language is to submitted. It's just like how this language grammatical and structural rules are constructed. I think that's the more important part. That's how many like base characters you have.

So it could be it could be that German is in fact a very efficient language because there's so much state in one word or one German efficient. I I would actually say for LM's German it may be efficient but it also may be inefficient because yeah if you have all of these complex words which which are like unique but then the LM also has to learn all of that and you have way less training data on each word. So the question would be like how specific can you go before the LLM has too little context and just breaks apart. Um, do LLMs face the same issue that we face with words with verbs in German that are um, gender?

No, that that they are in they it's it's a verb made of two words or like two half words and it can get separated. So, it's pretty it's pretty difficult to wrap your hand heads around it. that you start the the sentence by h by using half of the verb and then you use the other half of the verb at the very end of the sentence and it's a long sentence. You get to the end of the sentence and you get to the end verb and you ask wait what was the base of the verb? So you and it's very separated. So I'm thinking like does something like this even matter for LLMs because um I is I don't I don't think the LLM will have much difficulty with that because basically during the training process it's just like it within a sentence it weighs like which are important which is not important and based on that it generates it understanding I don't think it matters much if it's like side by side or if they are spaced apart also if even if you let's say we have a variant which is threatened together and a variant which is separated what many tokenizer do they have a word with a space in so the variants without space and with space are two different tokens and in the mind of the LLM distinct words in that aspect.

Okay. If you want I can also we could also talk about LLM tokenizer and just play around and I can show you what I mean. Oh yeah you you showed me once. Um can you search for llama 3 tokenizer online? I can show them exactly what I mean with the space. This will take roughly a minute. Yeah it's actually pretty cool to see it. Um and we maybe we numeric representation as well. I'm doing it my end so we can share the screen. Yeah, you can. Okay. Yeah, you have it. Oh yeah. So basically write hello and then write hello together and then write it once of a space. So four tokens together. Yeah. Oh, still four tokens.

You see how the color of the tokens changes because it's the it's Oh, wait. Does it also show which numbers these tokens are? But let me check this website. No, let me open a different one. Uh search uh offline. These are numbers here. Is that what you're talking about? Yeah, these numbers. And if if you Yeah, show these numbers. And now add the space. You see how the number changes in the LM? It's two distinct different tokens. And that's also like efficiency. I mean the entire term hello is 9906. Even if we for us it's five letters. Hi G E L L O and that's what I mean with like it depends on the tokenizer how efficient the language is.

Yeah. Yeah. Very cool. Right. Let's go to a break and uh we'll get back with some clawed code agents. Support for this episode comes from Flowbase. If you are building in Web Flow, Framer or Figma, Flowbase can help you build much faster. With over 4,000 components to choose from, you have a huge variety of wireframes and super clean, nicely designed sections that you can put together by just simply copy pasting them to your project. Over 15,000 icons as well, over a,000 illustrations. They also have a super helpful Web Flow app called Boosters. With boosters, you can add things like sliders and mares and countups to your project with a few clicks and without any coding.

So yeah, big thanks to Flowbase for supporting the show. If you want to try them out, go head to flowbase.co. And just for you guys, I want to show discount command flowbase all capitals for 20% off any plan. This is a limited offer, so get in there early. All right, that's it. Let's continue with the show. All right. So, clawed code sub aents, which I haven't used yet, but I'm particularly excited about. And these are basically preconfigured AI personalities that you can delegate task to inside of Clawude Code. And each kind of um each I guess agent has a spec specific expertise area, which is really cool. Use his own context window. It's separate from the kind of main conversation.

can be configured using tools. If we have a look down here, we'll see the kind of structure here. You can give it access to specific tools or not if you don't want to. Um, and you can like have it like do certain things and and um can configured uh with a custom system prompt. So, it just does really interesting stuff like with the prompt itself. Um, you can configure it as well to be like in the project here in the project or as your user. So you might have one that you constantly run as a user but then also ones that are very specific to a project. And it's just a really and you can chain sub agents together as well.

So, this is a really really cool um addition to Clawude Code, which I think is basically just the leader of uh of coding right now. It's my favorite. I I I use it. Um so now your AI agent has its own assistance. So if you are not failing enough, your AI has enough agents to fail double or triple. If you are not writing writing enough shitty code, yeah, more I actually disagree. I think this would improve stuff. Yeah, I like jokingly. I think this is pretty cool. Especially if they can check each other's work. Can they be critical of each other's work? getting they can spin up their own sub agents, but I don't think they they have their own context window.

So they're they're kind of independent of the of Well, it says independent of the main conversation. So yeah, so this is a code this is some example sub agents. So you've got a code review. So you can set up a code review agent to kind of just go off and do code reviews. So with every task, say you do um a feature, then you do a have a code review agent, you've got a debugger agent. Um well, they got a data scientist one here. Um and you can run all of these agents off as a I guess a consequence of the thing that you've just done. So it always makes sure it does an accessibility scan or um the most interesting thing is you can obviously if you've got something like playright you can feed the design sub agent you can feed in the image of the of the page it can spin all that up and whatever it's uh and to do all that just with again as a side effect of just one prompt is really cool they talk about like the best practice here is that they you start you start building with claude so basically you run uh slash agents and then you tell it what it wants to do.

So that's just a quick way to basically create the files which by the way as with everything with AI they're all markdown files. So once it's created the file the folder structure you can kind of go in then and then start adding you know you could create the first sentence but then you can go in and start adding a bunch of other stuff here. So yeah, start with the claws um generated uh design focus sub aents create uh sub aents with single clear responsibilities rather than trying to make one sub aent do everything. So they're really trying to isolate your your individual sub aents um to one specific task, right? Detailed prompts including instructions, examples and constraints, the normal kind of stuff you would probably give uh an AI for context and things like that.

limit tool access, which I really like that you can limit the tool access. Again, you could say the design sub a uh sub agent only has access to playright and only has access to magic UI or something like that. Um, and yeah, check project sub aents in version control, so you can uh collaborate and stuff like that. But really, and and here we go. Here's the chaining of sub aents as well as it talks about. So, it's a really cool um uh addition to to Craw and I'm just I just think it's getting better and better. I I like the idea. The more I think about it, the more it makes sense. Oh, yeah. The thing is I'm surprised they didn't have it yet because that's fundamentally how we will go on about L&Ms.

L&Ms as we've seen with like fine tuning everything. We can even fine tune a very small model to be better than large models at specific task. I think the end goal literally is and later on when fine tuning is easier we have literally LLM fine tuned on specific libraries specific to it and you just orchestration of first a compos and then let's say manage LLMs a little luck at the company would then literally build it feature for feature version for version and I think the next step here actually is to even have LLM spin up like duplicate directory and have each LLM work on a different feature and then merge it again with like merge request and have them work isolatedly.

So I think this is the intermediate step we reached now to where we truly have like huge swarms of LLM agents working on multiple features concally and pushing them out quickly. I mean we've had variants of this. I think I've seen hacks that people have managed to get Claude code to spin up its own sub agents and whatever like people have hacked it to be able to do that. But we've also got things like um codeex by openai or jewels by Google or I think GitHub's got one where they spin up agents in their own VM which is it complete contextual has it in theory has its own as you were saying its own um instance of an LLM you know maybe you can define what LLM it's using like a pre a pre-trained um LLM or something like that.

So they're they're kind of like we we're sort of we are figuring out like ways to like work. But I like I do like this idea and especially as um LLMs get cheaper faster and all the rest of it that we can start to see what you're talking about which is having dedicated LLMs that are just kickass at copywriting or or something like that. Um but this is just built in now and there's no need to hack the system. There's um and there there is like a minor hit on performance on this like there you know obviously without with anything you you start to like you know hit performance limits especially as you're churning through those GPUs and forest forests that you uh have to do um because yeah a few weeks ago they had a down like they uh anthropic like literally was like I was getting errors because they were just hitting those performance kind of things which is probably in part because people were spinning off sub agents and and whatever.

Um, saying that as well, there's been some pricing updates to Claude where they are setting weekly limits. So, cuz again, just people are um hammering their servers. Now, I watched a video today which kind of like made me feel a little bit better, but it it was just a small channel, but he made a video on why you why you don't need Claude Max. And the long and short of it is is less is more is a thing. I I'll I'll link it below because it's only a small channel, so might as well support, you know, where we can. But like um I have the pro plan. I don't pay the $100. I don't pay the $200.

I pay the $17 a month. Yeah. Granted, I don't code every single day. Like I'm not I'm not using it that much. So there is that aspect but also I understand the I'm not relying totally on claw code to write absolutely anything and in fact sometimes the plan I'll ask it to plan and I will implement the plan and you know my knowledge of code is is a lot better and I see a lot of comments on loads of coding videos complaining about you know uh things going wrong and this and that and and I can't help but feel like they are so dependent on the code on the AI doing all of the work for them.

They they that's all they've got to blame is the is the is the LLM uh the you know claude or whatever it is that they're using that things are going wrong and they're churning through their weekly limit in like no time at all. And it's like it's because you're it's because the these tools are a crutch to you. You're not learning anything and you're probably writing really inefficient prompts as well. But all that to be said is like yeah, I'm I'm I'm chilling through the lowest tier um plan of of Claude and loving it. Like I really like Claude. It's my favorite model to to use for coding. And then we got I also we also had the study regarding AI with the brain and that it makes let's say did you guys talk about that with the study where it says limps make you dumber?

Yes, we did. Yeah. Yeah. Yeah. Yeah. Um I saw an analysis of that and the interesting thing is um LM generally make you dumber if you don't know anything about the topic but let's say you're well where well where well where well where well where well where wellwhere well where well where well wherewell where we got a programmer if you use LLMs it actually even boosts your learning rate and the thing is what this generally state teaches us is you need to understand what you're doing because an LLM won't save you but once you understand what you're doing and you use an LLM it's going to boost your productivity by a fair margin be it learning and be it actually creating software also s one interesting thing I heard from you use the LLM to plan the edge texture and then implement it yourself I actually do it the other way around I plan the edge detector I literally tell it which feature to do it which endpoints to write in detail and I just let the LLM implement that and that's an interesting thing on how people it well no it's not something I do all the time but that's just that's just sometimes if I'm feeling like you know it's like driving a manual car instead of a automatic car it's like if I feel if I'm feeling like I want to code.

But no, I think the the way to what we're trying what we're settling on in in a from a coding perspective is you define your PRD, right? You define the actual structure of the function. You put a lot of time and energy into defining that and giving that to the LLM as opposed to just add this feature, you know, just a simple prompt to add the feature. So, no, I I absolutely do both. And I'm really experimenting with this approach of like how much context do you need to give it um to or or how much definition do I need to give it for it to do a good job or how much can I backst step because realistically everyone would want to take a complete backstep and just have the LLM figure it out for you.

But like we're definitely not there yet. And I think um yeah, a bit of a a nudge is is necessary. I'll do that method if I'm saying how can I improve this file or what what give me three examples or three ways I can I can refine this function or something like that and then you know copy paste or or implement it or whatever. It's just it's more of a brainstorming um you know soundboarding exercise as opposed to something I would always do. Um, so, so, so if you tell an LLM, so if you tell an AI something very generic, you become dumber. But if you say something very specific, you can actually learn from it.

And every time you are too generic with your prompt, uh, you also get something like really low quality. Like if you say make something beautiful, it doesn't really do a great job. It will give you something at best average. But once you know and you understand what is something beautiful or what is a great code architecture, you can hint at those specifics or literally tell it to to apply those specifics and then it can do them and you can actually even ask for exploration. So one step uh above this would be to to mention those specifics but also ask for uh explorations or alternatives but because you understand the basics um then you are talking in a whole different level with the LLMs and you also get information that you might not had before in opposed to be very generic and just ask ask it to build something to write something very generic that you don't understand.

So be if you are specific and you are also trying to learn from these specity uh you can actually learn from LLMs. I um I was I've actually been thinking about this dumb thing today actually and I was thinking I I think I read a comment somewhere or whatever. Um, someone did a review of uh, probably coral code because I watch loads of claw code stuff, but someone suggested using your voice because I think a lot of people are using their voice to to write their prompts to to whatever. I actually find it quite hard, but that's neither here nor there. Um, and someone said, you said to use your voice, but I don't know.

I don't where where do you turn where is that feature? Right. and they were looking their brain is completely switched off that they're looking for the the feature to be given to them when I translate that as well every single computer every single phone has a dictation button like you don't need to build a voice function to like type out a message unless they meant like an interactive like response and and whatever mode like a whatever I don't know what you call them conversational speech conversation mode yeah every every uh you know device has some sort of dictation. So, that's that's what triggered this thought for me of like why are people um switching off when it comes to stuff like um I think I sent a tweet out months ago that said something along the lines of people aren't I think it was someone looking for some sort of editing software or or gra like some sort of um uh a visual effects tool or something like that and someone was looking we we we we're kind of like being um conditioned ed to look for a SAS tool or a SAS product to do the exact thing we want.

We we've lost the ability to sort of think well broken down how can I achieve that with After Effects with this or with that right and it's because of convenience. It's because we're being given this convenience aspect and our brains just don't need to think. So we inevitably get dumber and and what is being dumber? I don't know. Maybe you just maybe it's maybe it's lazy. Maybe it's not being dumb. Maybe it's lazy. I don't know that. Um but all all that to be said there is the there is the um other aspect that um we are just well we just want we're just leveraging AI for what it is but you will always get people now design is an easy example or even development now development is so easy we're getting an influx of developers who have no ambition to learn how to code and we're we're like there like observing these people coming and being like, "Oh, they're so lazy and this and that."

It's because it's been democratized and we're just getting people, you're always going to get people who are lazy or dumb or whatever, um, coming into your industry and it's easy for us to stand on our high horse and think like, "Oh, they're just stupid or or or this or that." All that to be said, you're always going to get people who just who who don't let the tool make them stupid. They actually, as you say, allow it to make them smarter. They use it to learn. They use it to get better. So, I think we're just we're just it goes back to this idea of we're just a hyperconnected world where we're now able to see people or we're able to watch a video footage of someone being super stupid when they always existed.

You know, you're always going to get these people who are lazy and and dumb using the same sort of tools that we use. But, um yeah, I think it just comes down to just being super hyperconnected. So, who are we getting dumber? I just think that we're just being exposed to people who have always been there. They're just now entering our industry that we know a lot of and we can all we can do is sit back and think that they're just being stupid. Yeah. What we can definitely say is someone who was definitely lazy were the developers of the T app. Oh yes. Wow. Good segue. Good segue. Good one. Good one. Because uh what Yeah.

A very short summary. A T app is basically an app where um girls could rate images of guys and write comments and give like green flags, red flags. I'm not 100% certain, but that's something what I roughly what I heard. And basically what they did is in Firebase Oopsy. And that obsy leaked um 72,000 information about people including messages and images and location and location pretty much everything I think IDs as well. And it what's also really yeah it's really bad because this app was made with the mission to uh help women or like protect women. It's essentially about rating bad guys in a way or like be aware of bad guys in the dating world and talk about them and you you know people have different opinions about it but it was made for women to you know to to help them in that way but it actually hurt them.

Yeah. Turned it into a target. Yeah. How did that happen? So apparently, so it's important to say that this is well, first of all, the uh the uh CEO is a self-proclaimed uh C CEO of a digital company who's only done six months of, you know, coding boot camp, right? So he wore that with pride on his LinkedIn profile. But this is this is a breach that happened on all accounts that existed before February 2024. So any accounts that were created after that aren't affected by this breach. But ultimately it's the it was two there was actually two breaches. There was it was a two two-step program. The first one was that they found all of just the images the the um the verification images of the people.

So when you sign up you have to take a selfie of yourself holding your ID and that verifies that you're actually a girl and that you're actually a human. Um, and I mean first of all there were it was promised that those images would be deleted but clearly not. Um, but that's neither here nor there. That someone found the it was a public bucket stored on Firebase and some and because of the way the architecture of Firebase you're able it's kind of like an API and storage system built in. Normally on an app you separate those. You have an API layer between the two. you never hit the you never hit the uh database directly.

But with Firebase, it allows you to hit the database directly and if you hit a certain URL, it will list out all of the items in that bucket, which is something you can't do or you shouldn't be able to do anyway. Um, so what's someone done? They found the bucket and then they hit the they found an image, then reverse engineered that to be able to hit the bucket. They got their list of images and then just downloaded all those images. Um so that was the kind of first step of it and then someone found the actual um information on these people and there was there was a few apps created. There was a a location app so they pinned everyone's location in North America to find out where they all live.

Um, and there was the funniest one, which you know, I don't know how I feel about it, but there was is a website ranking these women based on how they look and, you know, using internet is cool, man. Very cool. Very cool. Um, and then there was a demographic one. someone I even got a uh I've even got the the tweet here like someone put some demographics um of the people of the women that they uh found. You know what this means is I'm not too sure. But the point is it's that the internet had a lot of fun. As cruel as it was, the internet had a lot of fun with these. Um it was a bit of a lame apology saying I've got the apology here as well.

um from the co saying a big company or is it just that they happen to have tons of users? Um I don't know the size of the company to be honest. Um it's not like it was um yeah I don't know the size of the company. No idea. But it was doing okay. I mean like I mean they had a wait list but this is like 15,000 likes. Yeah, I know. But like you you say that as if like I mean this is now they've been exposed, right? So they're probably a lot bigger now than they were to be honest. Yeah, but I mean like people are liking the apology meaning that Oh, I see.

Well, this is the thing and we won't go more into it, but this is the thing with the downvote button when on YouTube. What are they downvoting? Are they downvoting the or even up voting? Are they are they downvoting the idea of the video? The video is perfect actually. It was well presented this. But are they downvoting the idea of the video that the creators have nothing to do? You just don't know. So you can only speculate that they are they don't necessarily support the apology. Maybe they just find it funny. So they like it like people whatever. But the most Yeah. Sorry. It looks loweffort the image. This is their vibe though. This is their vi like I know what you mean.

I I I would tend to agree, but you can tell their design is a bit brash or and whatever. But this is the interesting thing. This data was stored to meet law enforcement standards around cyber bullying prevention. It doesn't say anything about what's the what's the ISO standard? It doesn't say anything about like data protection or anything like that. So yes, maybe I don't know what how you'd store an image to prevent cyber bullying. Maybe it's maybe that maybe that is simply that another user cannot access your imagery through their account. Maybe that's what that means. But cyber bullying prevention isn't a, you know, wellestablished standard in the in the um internet world which prevents the hacking of data.

You know, that's why the ISO standards are in place. So yeah, it's um a bit of a cluster to say the very least. And it's there was speculation as well that it was a vibe coded app obviously because it's like um in in the midst of all this vibe coding like slot that's being put out, but again because this is pre2024 like February 2024 um users that it's this definitely wouldn't have been a vibe code of that. I want to try and find the the CEO's name. Um, and while you add that, I could uh point a different topic. This is I would not say this is a specific issue to te. I think this is a general issue with Firebase.

Um, made once a video about Eva. She's a like a umi cyber like cyber security specialist and she once run a scraper through the entire web and found like six million affected Firebase in that sense with some kind of security issues where you could access the data and it's like people just don't know how to use Firebase and Firebase is like a gun which easily lets you shoot in a F. It's true like I've never liked Firebase. There was a lot of comments as well about the fact that Firebase pesters you because it's it's not trivial. Well, it is trivial, I guess, but like it warns you several times that this is like your this is wide open.

You haven't got security rules and things like that. It's it's not a great platform for proper app development. Like I say, it allows you to access the database directly, but there are security rules in place. And this was it last year that there was um because you have to put the API keys in the front end of your website. Yeah. And you have basically like security rules where you write a JavaScript like syntax. It's not JavaScript. It's just a JavaScript like syntax. And then you have to write security rules. And these can be a bit more comp complicated because you have to check I get the user ID and I have to check does you have access to this resource and it's like Yeah.

Just have a server. Yeah. Yeah. Yeah. It's not it's it I I never like those security rules. Here's a question I was wondering. Why is Superbase any different? Because I can access the database directly from the front end. I don't need to have some sort of API layer to query the database. Are you have you are you familiar with Superbase? Have you used Superbase? A bit. But the the reason why I like superbase more is generally from what I know I may be wrong on this but from what I know um you still can easily deploy functions with Firebase. If you want to deploy functions you have to then go to Google cloud and do everything like it generally doing things properly takes a bit of more work.

That's why I first like Firebase more. I don't use Firebase much. The second thing is um I if I recall correctly in Firebase you have more security like generally and it's like harder to shoot yourself in the foot. I may be wrong but from what I recall they generally have just more uh just better security. Um well I you can yes you could you've got the functions aspect but like again I can just query the database um and you use policies to like prevent you know and and Firebase has its own interpretation of policies and things like that but the point is is like I can access superbase in the same way that I can access uh I don't need to go in through any sort of API layer.

Um, saying that I can't find the CEO. Maybe he's maybe he's taken his uh account down or something. But way back, his name's C. His name's Cook. Uh, T CEO. He cooked himself. Did he Did he cook? He did. Yeah. Sean Cook. Sean Cook only has six months coding experience under his belt. Sean Cook. Let's um um what was I gonna say? Oh, saying that. Yeah, like I'm using I'm using um uh comet now and one thing I'm what I'm finding frustrated which I know we keep talking about like SEO is dead, but sometimes I literally just like it try sometimes because you've got the unified search bar. um I just want to find the website or I just want to go to the website or see like I know I know I want a list of website I don't want the summary so I'm having kind of like a reverse kind of like I know traffic is dying which is what you guys said last week here we go um but I'm also finding it sometimes more often than I care to admit I was weird um I'm actually wanting a traditional just give me a list of websites or give Like I don't want the summarization.

What the hell is going on? Been there with DIA as well. Sometimes, not too often. Uh I give it spec especially if I give it a long URL instead of Google searching it. It starts to reason and try to understand it. No, I don't want what I usually do. I just use sh and I say, "Hey, make me a message about this topic and give me a list with with a list your sources." And then it always says after each like link the source and you can just click it and go to the page. So if you that's what I've done. Yeah, that's what I do. I just click on sources now which is just an extra step because sometimes it's like uh I sort of know what it is I want and I'm expecting like a like a website.

Got to bloody sign in now. But um yeah, I wanna I want to try and get used to the kind of whole background uh thing like Sean Cook one. He is the first one. Um this guy. So maybe I think he's taken it down. I think he's taken it down. He had in his in his uh thing here, he said like six months of coding under his belt, which I he probably did not develop this app. Let's be clear. Like I just think that that was just being taken wildly out of context cuz it's whatever. But this is famous. Sorry. This is one way of becoming famous. Maybe it was clan. It was make a super bad app not secure at all and you'll be famous.

I mean there's tons of them out there and as as you said like the again last year there was a massive there was so many um breaches from just specifically the ruling around Firebase apps and the databases and object storage and stuff like that is just so clunky and easy to miss. Not easy to miss but just like you know most people just like leave open or whatever. So yeah um there's tons of apps out there. I think this was just I I honestly think there's a there's a touch of misogyny in in why this blew up as well because first of all it's like with the tech world we're ruled by guys so it was easy for us to you know what you know whether you have your own personal opinions on this type of app but there are don't don't sleep on the fact that there are so many apps being breached so many massive apps one massive one was very recent I forgotten what it was And it, you know, it might even be like a cyber security app or something.

But the point is, this happens an awful lot. So why this one's picked up a lot of steam, I think, is the misogyny aspect. Yeah. I mean, this is qu if somebody wants to ride on that wave, could make an app for the guys. And probably somebody already has or you know doing this. You know what I mean? I'm starting cloud code. Yeah, we call it instead of like tea, we call it social coffee. We call Well, isn't it Isn't the name is about like spill the tea? Like that's what you'd say is spill the tea. Yeah. Okay. So, yeah. You dropped the We dropped the handlebar. We dropped the handlebar. Yeah. Do you know I haven't been keeping an eye on um comments, so I'm just trying to Well, you can make an app uh to do the rating just exactly the same, but it's for us.

Yeah, it it's for for men. Lift up or lift down. But but for real, like like on waves like this, you could easily, you know, ma make it into news. Oh, yeah. No, it's um it's uh it's a catastrophe to be honest. But that was a that was a fun that was a funny segue because all we've got now is cursor to talk about and I was even considering removing it because it was a little bit boring. But we do have a new version of cursor that's been released. Um shared agent terminal uh uh context usage statistics which is something client code always had. I'm I'm happy to see it in I'm trying to get the swear thread.

Uh which is good here. So you can understand how many tokens you're using and the context size and stuff like that. Um and speed improvements to agent. So that's that's cursor 1.3. Nothing too interesting there. Yeah. But that is the week's news this week. Not too many thing h happened in the AI world. Well, we are a day early and true to form, I bet a bunch of stuff will happen tonight or probably happened whilst we were on this uh on this stream. Um because I thought it started off really quiet and I was like, "Oh man, we've got like an earlier thing and nothing's happening and then all of a sudden um you know the browser starts which I thought was quite good because obviously we had Max on earlier from Yeah.

from Warman. I would have hoped that we would have had more chance to speak to him about sort of like certainly from a security aspect as well. We could have probably gone into the security aspects whil talking about tea and stuff of of having and being signed into warm wind on you know a far far away machine somewhere. Um but as it is these people are very busy so good. Yeah. What would interest me really with that app is concurrency because let's say they advertise like customer support and everything but how do you manage which agents actually which open which email because if they if I have five agents and they all logged in into email and the mail comes in which agent takes over do they do you need to have some kind of scheduleuler and you need to somehow specialize it for an application.

Imagine I have agents and I tell them hey build me a web flow website and now imagine they both open the same project and try to work concurrently and then they try to modify the same component. It's like a lot of these things can happen and you need to some you also need to have some kind of orchestrate the agents so they don't run into each other and even in companies it's difficult. How do you do it with let's say highspeed agents who try to touch everything because LM are touchy these days. Yeah. Yeah. Interesting. Yeah. But speaking of web flow web flow was down this week uh for well Oh my god. Yeah. one week I actually get hired the one week I get hired to actually do something in web flow it's just been I thought it was just me I thought oh I' a new machine and this and that but it was it's been really really bad yeah so we had um a few not just one apologies from the CEO so web flow is acknowledging this and it's apparently uh one of their data providers and they specifically said it's not because of uh making new features.

It's not caused by new features. It's just a data. They don't know what it is though. They don't Well, they said they are Yeah. investigating and they are back online mostly, but they are still looking. Um, but I think they generally know what it is not. For example, it's not because of the new features that they are adding. Yeah. because people speculated it was the well, it was actually Christian specifically speculated that it might have been the the multiplayer um feature they added. Um, you know, I noticed it as soon as I started playing with Web Flow Cloud and this I was like, is it is it a Web Flow Cloud thing? Because I'm I literally, no word of a lie, probably spent about four hours, maybe even more, recording a tutorial yesterday because I just had to wait and builds kept failing and like back end wouldn't load and stuff like that.

I eventually got it recorded, but um yeah, it was it was very painful. And then I had a job that I actually had to do, but they fixed it in time for this job actually to be fair. But yeah, apparently Framer had an incident as well. So it wasn't just web flow. Oh wow. Yeah, probably it wasn't that big. Um it didn't get, you know, talked about too much, but framer had an incident as well. Interesting. and and and and Web Flow mentioned about the the sort of backend partner. I think they're just on AWS. I I can I can for sure guarantee Well, maybe not. I can't guarantee you, but I'm pretty sure it's not AWS.

Yeah, but what I know they use AWS, they do use AWS, but they also use other tools. They have tons of other services, and if you know, one in the chain doesn't work, they get into these troubles. But that was the news for Web Flow or like the big news uh of the week. And what I found interesting, people were really frustrated and very emotional about it. I get it. It happened to me as well, but it was nothing so that I couldn't open something. But I I did have I I did run into some loading issues, but not to the point to like I saw on Twitter, people were quite emotional about it. you're you're at the end of the day you're paying you're paying a premium for a service and like people do like I don't know I I definitely see both side I saw Joseph Barry saying you know we should be you know everyone has their bad days and whilst you know that's a really nice admirable mindset to go um into this stuff with or to calm yourself down at the end of the day this is they're literally costing you and or losing you money and you pay you pay for this sort of reliability, you know, resilience of an app.

So, I I just sort of see both sides of it. And um you know, it's one of the reasons I think I've made a video on this. I've seen Rand make a video on this. One of the reasons why you wouldn't choose Web Flow is because if something goes wrong with Web Flow, your business is down. Like, this was the thing that you signed up for. If you did your due diligence before you signed up to Web Flow, you would have watched one of our videos and we would have told you you are signing off a lot of your responsibility or your the reliability of you being able to do your work to web flow and they let you down.

The thing with that is in my mind it's a argument and I'm a developer who doesn't deploy anything web. I just got everything on my own and host my own VPS. The thing is at the end of the day you're always reliable on something. If you don't have AWS, you're reliable on that. If you have no VPS, you're reliable on that. You reliable on that the operating system update correctly that the browser works that your internet provider works. There's so many dependencies everywhere and it's web flow is badly said just one more way you consider where you like combine a couple of providers into one. I mean if web goes shitty but if your VPS provider goes on you have the same kind of issue.

It's like Yeah. I mean, if you were on WordPress and it was last year, well, was it last year? Yeah, last year. We We haven't forgotten that yet. And do you know what? I uh because they overtook WordPress overtook um advanced custom fields, right? Just took it. Just took it. Yeah. And I actually sw I was like, you know what? I use advanced custom fields. I'm going to stay on advanced custom fields. And then I found out the repeat literally just this week found out the repeater which is a paid for feature on advanced secure custom fields which is what they call it now it's free on there so I'll just update it to secure custom fields it's like we're so like if something's free or whatever like we'll be easily convinced otherwise but yeah so there is drama and vulnerability in any tool that you use across the internet no matter where it's But with web flow it's a case of honestly I I I would say it's bad luck.

I don't think that they are they've been stable. Yeah. Okay. They had sometimes connectivity issues they have been mostly stable like they have very high up time but the development tools been occasionally instable but the uptime is important that it's just always accessible. Yeah. I mean even even WhatsApp and Facebook they they have billions of users they have downtimes like I don't think people go on sorry I don't think I know of a single tool in my lifetime which never had an issue. Yeah. Yeah. I think it's because people's businesses operate from web flow. That's the problem. Businesses operate. People are earning money. Like it's not just what's I mean in theory your wife could be giving birth and you want to try and call her on WhatsApp but like you know well it's like also that you are paying you are paying for the service and there is also a service level that you know the 99% I I wonder how it was for the enterprise I don't know anything about that uh if the site was also down the dashboard was down for enterprise as well because that that That is where I assume web flow is making a lot of money and that is that would be really bad to hurt the relationship that they have with their client.

I have seen nothing about dedicated you know um servers for enterprise. No, but they have dedicated they have a different service level agreement with guaranteed SLA but that agreement might not. So the the typical uh website that you have is a 99.9% uptime guaranteed but with um enterprise I believe they add two more nines to the end. We also mix up top sorry up time. Yeah, because we we so that I stop to you but we're mixing two topics up. We have uptime of the website and up time of the editor. That's what I was about to say. Yeah. Yeah. And update of the editor was affected not from the website from what I know.

Yeah. The the websites were uh online but the editor being affected had one like really huge issue that was even if you could uh enter the designer and design something and change something you could not publish it. The publishing was also affected and that what from what I seen online what pissed people off because they were making the changes very difficultly with a a lot of difficulty they were making the changes but then they couldn't publish the changes um well it should be fine now a funny thing is like if you are a web flow user and hit that issue you theoretically could use a similar tool to web flow but uh with so many other tools if you hit some issue like this you cannot like switch tools as easily because of the proprietary technology.

Yes. I mean vendor login is the thing with everything. I mean with even be iOS and Android be let's say all of these website builders be let's say your vs code with your favorite extensions you're always locked into something be it with let's say US currency like there's always a lock in it's just like which system brings you the most advantages and if something goes down which is the advantages you have and if you're honest on the last note basically what the tools you use are in your hand and it's always your decision about how much money am I willing to spend versus comfort versus security and general risk and that's a calculation one has to do for themselves.

Yeah. Yeah. Yeah. Very true. Cool. All right. Let's wrap it up there. Unless you've got any more news items in your back pocket there, Kabaza. Um not no. This week was this week was building our Web Flow app. So Oh, how exciting. Is that why you're you're all together? Yeah. Yeah. is developing the web flow app and we should be ready to publish or like have our first version in a few days. Um and we have yeah we have like exciting things to share. We are building like cool components today. We had a call with u a 3D developer. So we'll we'll have like multiple cool things to to share in our app. You'll be the news next week.

Just be you. I'll be gone by then slightly. This was my last time I was here with him. I leave on Sunday. Oh, where are you going back to? Oh, B is calling. He also lives in Germany. So, yeah. Yeah, I guess I I guess that the accent is a dead giveaway. He he he lives uh like five years five years. Five hours away. Yeah, I used to drive it back home. Five years. My my brain is just in space and thinking of light years, you know. You need to get some budding after the stream. Yeah. No, we we need to finish the app. True. True. Wow. It was good to It was good to have you.

Two weeks in a row. Good. So, uh yeah, I appreciate you stepping in and giving your perspective. It's it was good to have someone who actually deals with the code side of LLMs and actually talk us through what a tool call is. Cool. All right, guys. Well, uh yeah, good luck with the launch and uh we'll speak to you soon.