← Back
Theo September 29, 2026 30m

OpenAI fights back

Not yet indexed — Search & Ask will be available once this episode finishes processing.

Read full transcript 23 segments
  1. They're back. They're back. Okay, let me catch you guys up quick. Okay, let me catch you guys up quick. Okay, let me catch you guys up quick. Last week Anthropic dropped a new model, Last week Anthropic dropped a new model, Last week Anthropic dropped a new model, Opus 5.5, and it was unbelievably good. Opus 5.5, and it was unbelievably good. Opus 5.5, and it was unbelievably good. It was so unbelievably good that OpenAI It was so unbelievably good that OpenAI It was so unbelievably good that OpenAI rushed out two model drops, GPT-6 Soul rushed out two model drops, GPT-6 Soul rushed out two model drops, GPT-6 Soul and GPT-6 Luna. You might have noticed I and GPT-6 Luna. You might have noticed I and GPT-6 Luna. You might have noticed I didn't do a video on those models. didn't do a video on those models. didn't do a video on those models. There's a reason. They weren't that There's a reason. They weren't that There's a reason. They weren't that good. I was not particularly impressed good. I was not particularly impressed good. I was not particularly impressed with either of them and didn't really with either of them and didn't really with either of them and didn't really have much to say. But there was one have much to say. But there was one have much to say. But there was one other model I happened to get early other model I happened to get early other model I happened to get early access to that is now available for all. access to that is now available for all. access to that is now available for all. That model is called GPT-6.1 That model is called GPT-6.1 That model is called GPT-6.1 Soul. And this model made it very, very Soul. And this model made it very, very Soul. And this model made it very, very hard to film that Sonnet 5.5 video hard to film that Sonnet 5.5 video hard to film that Sonnet 5.5 video because I knew something was coming. because I knew something was coming. because I knew something was coming. Something surprisingly cheap, Something surprisingly cheap, Something surprisingly cheap, surprisingly capable, and most surprisingly capable, and most surprisingly capable, and most surprisingly better than Astra. So, surprisingly better than Astra. So, surprisingly better than Astra. So, yeah. I got a lot to say about this one. yeah. I got a lot to say about this one. yeah. I got a lot to say about this one. I'm filming this at 2:00 in the morning I'm filming this at 2:00 in the morning I'm filming this at 2:00 in the morning right after filming my Sonnet 5.5 right after filming my Sonnet 5.5 right after filming my Sonnet 5.5 videos. So, pardon me for stumbling over videos. So, pardon me for stumbling over videos. So, pardon me for stumbling over a few words here and there. I'm doing my a few words here and there. I'm doing my a few words here and there. I'm doing my best to get this out as reasonably best to get this out as reasonably best to get this out as reasonably quickly as possible because I want to quickly as possible because I want to quickly as possible because I want to have some coverage. And I'll be real, it have some coverage. And I'll be real, it have some coverage. And I'll be real, it is also quite fun to cover these things is also quite fun to cover these things is also quite fun to cover these things before they are out. So, you are getting before they are out. So, you are getting before they are out. So, you are getting my true honest take and not the my true honest take and not the my true honest take and not the distilled version of what everyone else distilled version of what everyone else distilled version of what everyone else is saying. I'm sure this model's going is saying. I'm sure this model's going is saying. I'm sure this model's going to cause some pretty crazy waves. So, to cause some pretty crazy waves. So, to cause some pretty crazy waves. So, it'll be nice to have my take out it'll be nice to have my take out it'll be nice to have my take out initially separately first. As always, I initially separately first. As always, I initially separately first. As always, I feel obligated to remind you I do have feel obligated to remind you I do have feel obligated to remind you I do have early access, but I'm not being paid in early access, but I'm not being paid in early access, but I'm not being paid in any way, shape, or form. OpenAI has no any way, shape, or form. OpenAI has no any way, shape, or form. OpenAI has no influence over what and how I say influence over what and how I say influence over what and how I say things, just when. They politely asked things, just when. They politely asked things, just when. They politely asked me to wait until the model is out to me to wait until the model is out to me to wait until the model is out to talk about it, which makes a lot of talk about it, which makes a lot of talk about it, which makes a lot of sense. But I have to wait for one other

  2. sense. But I have to wait for one other sense. But I have to wait for one other thing first. Today's sponsor. In order thing first. Today's sponsor. In order thing first. Today's sponsor. In order to build good software with agents, you to build good software with agents, you to build good software with agents, you need to get feedback. And let's be real, need to get feedback. And let's be real, need to get feedback. And let's be real, they're getting a lot of that feedback they're getting a lot of that feedback they're getting a lot of that feedback from our CI. That's why we've all been from our CI. That's why we've all been from our CI. That's why we've all been seeing our CI bills skyrocket. And also seeing our CI bills skyrocket. And also seeing our CI bills skyrocket. And also why we've been getting more and more why we've been getting more and more why we've been getting more and more frustrated with GitHub Actions. Today's frustrated with GitHub Actions. Today's frustrated with GitHub Actions. Today's sponsor is Depot, and they're here to sponsor is Depot, and they're here to sponsor is Depot, and they're here to solve all of this and more. Not only can solve all of this and more. Not only can solve all of this and more. Not only can they make your CI up to 10 times faster, they make your CI up to 10 times faster, they make your CI up to 10 times faster, as well as your Docker builds up to 40 as well as your Docker builds up to 40 as well as your Docker builds up to 40 times faster, especially when they're times faster, especially when they're times faster, especially when they're downloading cache. They're also cheaper downloading cache. They're also cheaper downloading cache. They're also cheaper and they give better feedback for your and they give better feedback for your and they give better feedback for your agents. All this is possible due to agents. All this is possible due to agents. All this is possible due to Depot metal. They're running their own Depot metal. They're running their own Depot metal. They're running their own bare metal with AMD epic processors that bare metal with AMD epic processors that bare metal with AMD epic processors that are way faster than what you get from are way faster than what you get from are way faster than what you get from traditional CI providers like, of traditional CI providers like, of traditional CI providers like, of course, GitHub Actions. If you wanted to course, GitHub Actions. If you wanted to course, GitHub Actions. If you wanted to be a drop-in replacement, it absolutely be a drop-in replacement, it absolutely be a drop-in replacement, it absolutely can be, but Depot's APIs are so much can be, but Depot's APIs are so much can be, but Depot's APIs are so much better that you should probably use better that you should probably use better that you should probably use those instead. They enable those instead. They enable those instead. They enable parallelization and most importantly, parallelization and most importantly, parallelization and most importantly, resilience when GitHub inevitably goes resilience when GitHub inevitably goes resilience when GitHub inevitably goes down randomly for no good reason. We've down randomly for no good reason. We've down randomly for no good reason. We've had our releases get blocked because we had our releases get blocked because we had our releases get blocked because we weren't using Depot and I'm so thankful weren't using Depot and I'm so thankful weren't using Depot and I'm so thankful that I've been moving more and more that I've been moving more and more that I've been moving more and more stuff over. For example, when Ben moved stuff over. For example, when Ben moved stuff over. For example, when Ben moved Pick Thing over to Bun, we immediately Pick Thing over to Bun, we immediately Pick Thing over to Bun, we immediately had some CI failures. Normally, this had some CI failures. Normally, this had some CI failures. Normally, this would be obscure piles of text that our would be obscure piles of text that our would be obscure piles of text that our agents parse through for us, but when we agents parse through for us, but when we agents parse through for us, but when we use Depot, it becomes way easier to see.

  3. use Depot, it becomes way easier to see. use Depot, it becomes way easier to see. They'll even analyze the failures and They'll even analyze the failures and They'll even analyze the failures and give suggestions to make it much simpler give suggestions to make it much simpler give suggestions to make it much simpler to get this feedback back to our agents. to get this feedback back to our agents. to get this feedback back to our agents. This is especially useful when you tell This is especially useful when you tell This is especially useful when you tell your agents that you can use Depot your agents that you can use Depot your agents that you can use Depot because they'll no longer have to push because they'll no longer have to push because they'll no longer have to push changes and wait for that to trigger a changes and wait for that to trigger a changes and wait for that to trigger a build. They can just run the CLI to build. They can just run the CLI to build. They can just run the CLI to trigger the exact same CI that you'd be trigger the exact same CI that you'd be trigger the exact same CI that you'd be triggering through GitHub instead. No triggering through GitHub instead. No triggering through GitHub instead. No longer do you have to file PRs of broken longer do you have to file PRs of broken longer do you have to file PRs of broken code just to get feedback to your code just to get feedback to your code just to get feedback to your agents. They could just run a tool agents. They could just run a tool agents. They could just run a tool instead. Your agents will also get way instead. Your agents will also get way instead. Your agents will also get way better breakdowns of what is taking so better breakdowns of what is taking so better breakdowns of what is taking so long in your actual CI runs so that you long in your actual CI runs so that you long in your actual CI runs so that you can figure out how to improve them and can figure out how to improve them and can figure out how to improve them and make them faster and more reliable. You make them faster and more reliable. You make them faster and more reliable. You and your agents deserve faster Docker, and your agents deserve faster Docker, and your agents deserve faster Docker, faster build times, faster CI, better faster build times, faster CI, better faster build times, faster CI, better results, and ideally a cheaper price. results, and ideally a cheaper price. results, and ideally a cheaper price. Get all of that and more at Get all of that and more at Get all of that and more at slite.link/depot. slite.link/depot. slite.link/depot. Let's talk about this model a bit Let's talk about this model a bit Let's talk about this model a bit because it is not quite what I expected because it is not quite what I expected because it is not quite what I expected and it's probably not what you guys and it's probably not what you guys and it's probably not what you guys expected either. Especially when you expected either. Especially when you expected either. Especially when you consider that GPT-6 sold just came out consider that GPT-6 sold just came out consider that GPT-6 sold just came out like a week ago. It'll be a round a one like a week ago. It'll be a round a one like a week ago. It'll be a round a one week gap from 6.0 sold to 6.1 sold. I week gap from 6.0 sold to 6.1 sold. I week gap from 6.0 sold to 6.1 sold. I also want to disclose the numbers I'm also want to disclose the numbers I'm also want to disclose the numbers I'm currently showing on my screen are currently showing on my screen are currently showing on my screen are unlikely to be exactly accurate because unlikely to be exactly accurate because unlikely to be exactly accurate because I am running terminal bench for myself I am running terminal bench for myself I am running terminal bench for myself in the first two times I ran it, I in the first two times I ran it, I in the first two times I ran it, I screwed things up. The third one seems screwed things up. The third one seems screwed things up. The third one seems to be doing much better. I didn't run it to be doing much better. I didn't run it to be doing much better. I didn't run it on medium initially so there's a missing on medium initially so there's a missing on medium initially so there's a missing there, but low, high, x-high, and max, there, but low, high, x-high, and max, there, but low, high, x-high, and max, although the max run is incomplete, so although the max run is incomplete, so although the max run is incomplete, so I'm currently backfilling scores from I'm currently backfilling scores from I'm currently backfilling scores from x-high for the ones that max either got x-high for the ones that max either got x-high for the ones that max either got wrong or in previous run or didn't do wrong or in previous run or didn't do wrong or in previous run or didn't do yet cuz it takes like 8 plus hours in yet cuz it takes like 8 plus hours in yet cuz it takes like 8 plus hours in some of these tasks. I This bench is some of these tasks. I This bench is some of these tasks. I This bench is nuts. It's I'm more skeptical of nuts. It's I'm more skeptical of nuts. It's I'm more skeptical of benchmarks than ever now that I've been

  4. benchmarks than ever now that I've been benchmarks than ever now that I've been running a lot more of them myself in running a lot more of them myself in running a lot more of them myself in order to get the coverage I want to give order to get the coverage I want to give order to get the coverage I want to give here. For what it is worth, Terminal here. For what it is worth, Terminal here. For what it is worth, Terminal Bench 4 is state of the art score here, Bench 4 is state of the art score here, Bench 4 is state of the art score here, as is Deep SWE, although this one's as is Deep SWE, although this one's as is Deep SWE, although this one's weirder because it goes down on x-high weirder because it goes down on x-high weirder because it goes down on x-high and max and stays even on low and high, and max and stays even on low and high, and max and stays even on low and high, but those even low and high scores are but those even low and high scores are but those even low and high scores are scoring around what Astra did on high. scoring around what Astra did on high. scoring around what Astra did on high. The difference being it's doing it for The difference being it's doing it for The difference being it's doing it for comically cheaper. If you switch over to comically cheaper. If you switch over to comically cheaper. If you switch over to the log scale, you'll see what I mean. the log scale, you'll see what I mean. the log scale, you'll see what I mean. This model on low costs 21 cents versus This model on low costs 21 cents versus This model on low costs 21 cents versus Astra on low costing $1.46 Astra on low costing $1.46 Astra on low costing $1.46 and Opus 5.5 on max getting the same and Opus 5.5 on max getting the same and Opus 5.5 on max getting the same score for $14.65. score for $14.65. score for $14.65. While I will gladly admit that Deep SWE While I will gladly admit that Deep SWE While I will gladly admit that Deep SWE is far from a perfect measure of how is far from a perfect measure of how is far from a perfect measure of how good a model is at day-to-day code work, good a model is at day-to-day code work, good a model is at day-to-day code work, the fact that 6.1 Soul is scoring the the fact that 6.1 Soul is scoring the the fact that 6.1 Soul is scoring the same as Opus and is also 73x cheaper is same as Opus and is also 73x cheaper is same as Opus and is also 73x cheaper is at least worth noticing. Here's where at least worth noticing. Here's where at least worth noticing. Here's where I'm going to say some of the things that I'm going to say some of the things that I'm going to say some of the things that I probably shouldn't. Considering that I probably shouldn't. Considering that I probably shouldn't. Considering that GPT-6 Soul came out last week on Tuesday GPT-6 Soul came out last week on Tuesday GPT-6 Soul came out last week on Tuesday and this model's coming out this week on and this model's coming out this week on and this model's coming out this week on Tuesday, I think it's reasonable to Tuesday, I think it's reasonable to Tuesday, I think it's reasonable to infer that 6.1 Soul was not meant to be infer that 6.1 Soul was not meant to be infer that 6.1 Soul was not meant to be 6.1 Soul. There are things I'm not 6.1 Soul. There are things I'm not 6.1 Soul. There are things I'm not supposed to say and I'm definitely supposed to say and I'm definitely supposed to say and I'm definitely walking a thin line here by sharing it, walking a thin line here by sharing it, walking a thin line here by sharing it, so I hope this proves I'm not paid off so I hope this proves I'm not paid off so I hope this proves I'm not paid off by OpenAI because I'm about to give you by OpenAI because I'm about to give you by OpenAI because I'm about to give you guys info that I Yeah, it just let let guys info that I Yeah, it just let let guys info that I Yeah, it just let let me get through this. First and foremost, me get through this. First and foremost, me get through this. First and foremost, 6.1 is significantly smarter than GPT-6 6.1 is significantly smarter than GPT-6 6.1 is significantly smarter than GPT-6 Soul, thereby indicating this isn't just Soul, thereby indicating this isn't just Soul, thereby indicating this isn't just a dot one bump. There is something a dot one bump. There is something a dot one bump. There is something fundamentally different here. Next point

  5. fundamentally different here. Next point fundamentally different here. Next point is that it has meaningfully slower is that it has meaningfully slower is that it has meaningfully slower tokens per second. That tends to tokens per second. That tends to tokens per second. That tends to indicate the model is bigger. Hard to indicate the model is bigger. Hard to indicate the model is bigger. Hard to know for sure. Seems like this model know for sure. Seems like this model know for sure. Seems like this model might be different. Most importantly, we might be different. Most importantly, we might be different. Most importantly, we have Tibo's tweet. What am I referring have Tibo's tweet. What am I referring have Tibo's tweet. What am I referring to there? Well, right before I started to there? Well, right before I started to there? Well, right before I started filming, Tibo dropped quite a wall of filming, Tibo dropped quite a wall of filming, Tibo dropped quite a wall of text. The thing I want to emphasize here text. The thing I want to emphasize here text. The thing I want to emphasize here is first off that the Pro $200 is first off that the Pro $200 is first off that the Pro $200 subscription is back, but more subscription is back, but more subscription is back, but more importantly, importantly, importantly, more importantly is this sentence. They more importantly is this sentence. They more importantly is this sentence. They are changing how they calculate the are changing how they calculate the are changing how they calculate the usage in the sub. In effect, if you do usage in the sub. In effect, if you do usage in the sub. In effect, if you do the math, it will net out at half the the math, it will net out at half the the math, it will net out at half the dollar in API spend compared to the old dollar in API spend compared to the old dollar in API spend compared to the old Pro $200 plan. Why in the world would Pro $200 plan. Why in the world would Pro $200 plan. Why in the world would they do this? Especially right now where they do this? Especially right now where they do this? Especially right now where there is allegedly an internal code red there is allegedly an internal code red there is allegedly an internal code red because Opus 5.5 is so unbelievably good because Opus 5.5 is so unbelievably good because Opus 5.5 is so unbelievably good and has made the $200 Claude code sub and has made the $200 Claude code sub and has made the $200 Claude code sub such an unbelievable value. The only such an unbelievable value. The only such an unbelievable value. The only reason in the world Tibo would post this reason in the world Tibo would post this reason in the world Tibo would post this right now is a I don't know. Maybe a new right now is a I don't know. Maybe a new right now is a I don't know. Maybe a new model is coming where their margins model is coming where their margins model is coming where their margins aren't as good, so the ability to aren't as good, so the ability to aren't as good, so the ability to subsidize has gone down because subsidize has gone down because subsidize has gone down because remember, you can get 8 to 9,000 dollars remember, you can get 8 to 9,000 dollars remember, you can get 8 to 9,000 dollars of usage in a month on the $200 Claude of usage in a month on the $200 Claude of usage in a month on the $200 Claude code plan, and you can get over 12 grand code plan, and you can get over 12 grand code plan, and you can get over 12 grand on the $200 Codex plan. I did actually on the $200 Codex plan. I did actually on the $200 Codex plan. I did actually run a lot of numbers before this, and run a lot of numbers before this, and run a lot of numbers before this, and the amount you could get on Astra did go the amount you could get on Astra did go the amount you could get on Astra did go down slightly closer to like 8 grand or down slightly closer to like 8 grand or down slightly closer to like 8 grand or so. Hard to know for sure because they so. Hard to know for sure because they so. Hard to know for sure because they differ for everyone everywhere, and it's differ for everyone everywhere, and it's differ for everyone everywhere, and it's hard to log all of this stuff, but from hard to log all of this stuff, but from hard to log all of this stuff, but from my math, roughly 9 grand a month of my math, roughly 9 grand a month of my math, roughly 9 grand a month of usage, and that's where the price for usage, and that's where the price for usage, and that's where the price for this model comes in. This model is $2 this model comes in. This model is $2 this model comes in. This model is $2 per million input tokens and $10 per

  6. per million input tokens and $10 per per million input tokens and $10 per million out. This makes it way cheaper million out. This makes it way cheaper million out. This makes it way cheaper than 5.6 sole was at launch, half the than 5.6 sole was at launch, half the than 5.6 sole was at launch, half the price of 5.6 sole after discounts, and price of 5.6 sole after discounts, and price of 5.6 sole after discounts, and the same price as GPT-6 sole, 1/5 the the same price as GPT-6 sole, 1/5 the the same price as GPT-6 sole, 1/5 the price of Astra. However, this is not the price of Astra. However, this is not the price of Astra. However, this is not the whole story cuz cash reads matter, and whole story cuz cash reads matter, and whole story cuz cash reads matter, and the cash read price for this model is the cash read price for this model is the cash read price for this model is going to be 10 cents per mill in. That's going to be 10 cents per mill in. That's going to be 10 cents per mill in. That's a big deal. OpenAI has not changed cash a big deal. OpenAI has not changed cash a big deal. OpenAI has not changed cash read price as far as I know ever before. read price as far as I know ever before. read price as far as I know ever before. It's always been exactly 10% of the It's always been exactly 10% of the It's always been exactly 10% of the normal read price. That makes it a 90% normal read price. That makes it a 90% normal read price. That makes it a 90% discount, and now it's a 95% discount. discount, and now it's a 95% discount. discount, and now it's a 95% discount. That means they cut the cash read cost That means they cut the cash read cost That means they cut the cash read cost in half, massively reducing the cost for in half, massively reducing the cost for in half, massively reducing the cost for real-world agentic use, which to be real-world agentic use, which to be real-world agentic use, which to be clear is what we're using these for most clear is what we're using these for most clear is what we're using these for most of the time. of the time. of the time. So, this makes the model absurdly cheap So, this makes the model absurdly cheap So, this makes the model absurdly cheap for doing real-world code work. That for doing real-world code work. That for doing real-world code work. That also means that they are almost also means that they are almost also means that they are almost certainly cutting into their margins. certainly cutting into their margins. certainly cutting into their margins. Historically, these margins are rumored Historically, these margins are rumored Historically, these margins are rumored to be as high as 95%, to be as high as 95%, to be as high as 95%, like for every $10 you spend, they only like for every $10 you spend, they only like for every $10 you spend, they only have to spend 50 cents. And as crazy as have to spend 50 cents. And as crazy as have to spend 50 cents. And as crazy as that sounds, it makes a lot of sense, that sounds, it makes a lot of sense, that sounds, it makes a lot of sense, especially when you consider how especially when you consider how especially when you consider how expensive it is to make and train these expensive it is to make and train these expensive it is to make and train these models. But, that also gives them wiggle models. But, that also gives them wiggle models. But, that also gives them wiggle room to change things around a bit, room to change things around a bit, room to change things around a bit, which appears to be what's happening which appears to be what's happening which appears to be what's happening here. That also means that if they were here. That also means that if they were here. That also means that if they were to keep subsidizing the same level that to keep subsidizing the same level that to keep subsidizing the same level that they were on the subscriptions, that they were on the subscriptions, that they were on the subscriptions, that your electricity cost for your sub would your electricity cost for your sub would your electricity cost for your sub would be more than you're paying. So, I get be more than you're paying. So, I get be more than you're paying. So, I get why they have to change this. They've why they have to change this. They've why they have to change this. They've kind of just left the details out there kind of just left the details out there kind of just left the details out there for us to reverse-engineer, so for us to reverse-engineer, so for us to reverse-engineer, so take this as you will. 6.1 coming so take this as you will. 6.1 coming so take this as you will. 6.1 coming so fast seems to indicate it is not just a

  7. fast seems to indicate it is not just a fast seems to indicate it is not just a new snapshot of GPT-6. So, let's talk new snapshot of GPT-6. So, let's talk new snapshot of GPT-6. So, let's talk more about this model. more about this model. more about this model. As I was showing earlier, As I was showing earlier, As I was showing earlier, seems really good at terminal bench. seems really good at terminal bench. seems really good at terminal bench. Every time we refresh, the numbers Every time we refresh, the numbers Every time we refresh, the numbers change because new runs come in, and it change because new runs come in, and it change because new runs come in, and it looks like Max failed some things that looks like Max failed some things that looks like Max failed some things that XHi passed, which is why it just dropped XHi passed, which is why it just dropped XHi passed, which is why it just dropped a bit. But again, pretty much all of a bit. But again, pretty much all of a bit. But again, pretty much all of these, even high and XHi, are scoring these, even high and XHi, are scoring these, even high and XHi, are scoring higher than anything else ever has. And higher than anything else ever has. And higher than anything else ever has. And this is for me running this benchmark on this is for me running this benchmark on this is for me running this benchmark on random VMs on my network, so random VMs on my network, so random VMs on my network, so not the best suite to test against. I not the best suite to test against. I not the best suite to test against. I also had to drop three particular tasks also had to drop three particular tasks also had to drop three particular tasks from it because they expected an H100 to from it because they expected an H100 to from it because they expected an H100 to work against, which I make decent money. work against, which I make decent money. work against, which I make decent money. I don't make H100 money, okay? But none I don't make H100 money, okay? But none I don't make H100 money, okay? But none of this is real-world code work. So, of this is real-world code work. So, of this is real-world code work. So, let's talk a bit about that. Obviously, let's talk a bit about that. Obviously, let's talk a bit about that. Obviously, we'll have all the fun things like fish we'll have all the fun things like fish we'll have all the fun things like fish slop near the end, so stay tuned for slop near the end, so stay tuned for slop near the end, so stay tuned for that. But I just want to fixate a bit on that. But I just want to fixate a bit on that. But I just want to fixate a bit on the costs here because the most the costs here because the most the costs here because the most expensive run with 6.1 sole for me was expensive run with 6.1 sole for me was expensive run with 6.1 sole for me was about a dollar and 38 cents per task, about a dollar and 38 cents per task, about a dollar and 38 cents per task, and to cheapest run with Opus 5.5 was and to cheapest run with Opus 5.5 was and to cheapest run with Opus 5.5 was $5.12. $5.12. $5.12. That's a 4 to 5x gap from the cheapest That's a 4 to 5x gap from the cheapest That's a 4 to 5x gap from the cheapest Opus to the most expensive Soul. So, at Opus to the most expensive Soul. So, at Opus to the most expensive Soul. So, at this point, I would imagine you are this point, I would imagine you are this point, I would imagine you are hoping and praying this model is good hoping and praying this model is good hoping and praying this model is good and that it can actually replace Opus and that it can actually replace Opus and that it can actually replace Opus 5.5 for day-to-day work. And I promise 5.5 for day-to-day work. And I promise 5.5 for day-to-day work. And I promise we'll get some good answers to that in a we'll get some good answers to that in a we'll get some good answers to that in a bit. But first, we need to talk a bit bit. But first, we need to talk a bit bit. But first, we need to talk a bit about model behaviors here. Because this about model behaviors here. Because this about model behaviors here. Because this model is a part of the GPT-6 family, model is a part of the GPT-6 family, model is a part of the GPT-6 family, which means it has a behaviors that are which means it has a behaviors that are which means it has a behaviors that are worth talking about.

  8. worth talking about. worth talking about. I know I cite this diagram a lot, but I know I cite this diagram a lot, but I know I cite this diagram a lot, but there's a reason for it. The thing that there's a reason for it. The thing that there's a reason for it. The thing that made me so frustrated with GPT-6 Astra made me so frustrated with GPT-6 Astra made me so frustrated with GPT-6 Astra wasn't that it was less intelligent than wasn't that it was less intelligent than wasn't that it was less intelligent than the best models from Anthropic, because the best models from Anthropic, because the best models from Anthropic, because it was more intelligent than the best it was more intelligent than the best it was more intelligent than the best models from Anthropic, and I would argue models from Anthropic, and I would argue models from Anthropic, and I would argue in many ways still is. But there was a in many ways still is. But there was a in many ways still is. But there was a problem. It is also dumb. It is smart problem. It is also dumb. It is smart problem. It is also dumb. It is smart and dumb at the same time. GPT-6 Astra and dumb at the same time. GPT-6 Astra and dumb at the same time. GPT-6 Astra would just randomly spike into the would just randomly spike into the would just randomly spike into the dumbest I've seen a model do dumbest I've seen a model do dumbest I've seen a model do this year, even worse than like some of this year, even worse than like some of this year, even worse than like some of the small open weight models I play the small open weight models I play the small open weight models I play with. It's still so deeply frustrating with. It's still so deeply frustrating with. It's still so deeply frustrating that Astra does this, because on the that Astra does this, because on the that Astra does this, because on the other end, when it does well, it's other end, when it does well, it's other end, when it does well, it's unbelievable. But these spikes got to unbelievable. But these spikes got to unbelievable. But these spikes got to the point where I effectively churned. I the point where I effectively churned. I the point where I effectively churned. I was only using my Codex subs for was only using my Codex subs for was only using my Codex subs for computer use, and I ended up just computer use, and I ended up just computer use, and I ended up just leaning on the Fable 5.1 and obviously leaning on the Fable 5.1 and obviously leaning on the Fable 5.1 and obviously now Opus 5.5 for almost all of my now Opus 5.5 for almost all of my now Opus 5.5 for almost all of my day-to-day work. So, have they addressed day-to-day work. So, have they addressed day-to-day work. So, have they addressed the spikiness? Has GPT-6.1 Soul fixed the spikiness? Has GPT-6.1 Soul fixed the spikiness? Has GPT-6.1 Soul fixed the problems that I was so frustrated the problems that I was so frustrated the problems that I was so frustrated about with Astra? I would say mostly. about with Astra? I would say mostly. about with Astra? I would say mostly. Not entirely, but for the most part, Not entirely, but for the most part, Not entirely, but for the most part, yeah, this is a much better model. Its yeah, this is a much better model. Its yeah, this is a much better model. Its peaks are not as high. This is not the peaks are not as high. This is not the peaks are not as high. This is not the incredible revolutionary 3D capabilities incredible revolutionary 3D capabilities incredible revolutionary 3D capabilities that we saw with Astra. In fact, I would that we saw with Astra. In fact, I would that we saw with Astra. In fact, I would put it slightly below 5.5 Opus in most put it slightly below 5.5 Opus in most put it slightly below 5.5 Opus in most of those types of things. It is not as of those types of things. It is not as of those types of things. It is not as good at computer use as Astra, although good at computer use as Astra, although good at computer use as Astra, although it is close enough to the point where I it is close enough to the point where I it is close enough to the point where I have been happy using it for all of my have been happy using it for all of my have been happy using it for all of my day-to-day work. I actually had 6.1 Soul day-to-day work. I actually had 6.1 Soul day-to-day work. I actually had 6.1 Soul go through all of my emails and find go through all of my emails and find go through all of my emails and find invoices that I had forgotten to pay or invoices that I had forgotten to pay or invoices that I had forgotten to pay or was behind on, mostly like investing was behind on, mostly like investing was behind on, mostly like investing stuff, and set up new tabs in Chrome for

  9. stuff, and set up new tabs in Chrome for stuff, and set up new tabs in Chrome for every investment I needed to wire, fill every investment I needed to wire, fill every investment I needed to wire, fill out all the details for me, and just out all the details for me, and just out all the details for me, and just leave me to hit send. It didn't get a leave me to hit send. It didn't get a leave me to hit send. It didn't get a single thing wrong and called out single thing wrong and called out single thing wrong and called out additional stuff that I absolutely would additional stuff that I absolutely would additional stuff that I absolutely would have missed if I was doing this work have missed if I was doing this work have missed if I was doing this work myself. So, I'm literally trusting this myself. So, I'm literally trusting this myself. So, I'm literally trusting this model to wire money for me. It's model to wire money for me. It's model to wire money for me. It's trustworthy enough for that, and trustworthy enough for that, and trustworthy enough for that, and honestly, I don't know if I would have honestly, I don't know if I would have honestly, I don't know if I would have trusted Astra with that due to the trusted Astra with that due to the trusted Astra with that due to the spikiness. 6.1 Soul, much, much less spikiness. 6.1 Soul, much, much less spikiness. 6.1 Soul, much, much less spiky. From what I've heard from the spiky. From what I've heard from the spiky. From what I've heard from the other testers, they seem to agree with other testers, they seem to agree with other testers, they seem to agree with this analysis. I know for a fact that this analysis. I know for a fact that this analysis. I know for a fact that Julius and Ben, who have also been Julius and Ben, who have also been Julius and Ben, who have also been testing, have had a much better testing, have had a much better testing, have had a much better experience with this than Astra in terms experience with this than Astra in terms experience with this than Astra in terms of the spikiness. Julius called the of the spikiness. Julius called the of the spikiness. Julius called the model incredible. Ben called it model incredible. Ben called it model incredible. Ben called it incredibly boring. And I think that's incredibly boring. And I think that's incredibly boring. And I think that's the best place you can be for a model the best place you can be for a model the best place you can be for a model drop like this. But, as I had mentioned drop like this. But, as I had mentioned drop like this. But, as I had mentioned before, its peaks are not as impressive. before, its peaks are not as impressive. before, its peaks are not as impressive. While it does quality work the majority While it does quality work the majority While it does quality work the majority of the time, there are some tasks that of the time, there are some tasks that of the time, there are some tasks that are just at the edge of its capability are just at the edge of its capability are just at the edge of its capability that it will start to do weirder stuff that it will start to do weirder stuff that it will start to do weirder stuff on. For the most part, it's fine, but I on. For the most part, it's fine, but I on. For the most part, it's fine, but I I'm still reaching for Opus a decent I'm still reaching for Opus a decent I'm still reaching for Opus a decent bit. We'll talk more about the bit. We'll talk more about the bit. We'll talk more about the comparison later. I do default to this comparison later. I do default to this comparison later. I do default to this model for a bunch of stuff, though. model for a bunch of stuff, though. model for a bunch of stuff, though. First off, as I mentioned before, First off, as I mentioned before, First off, as I mentioned before, computer use. I can't wait for ultrafast computer use. I can't wait for ultrafast computer use. I can't wait for ultrafast to be like an actual thing you can use to be like an actual thing you can use to be like an actual thing you can use with OpenAI models, because when it is, with OpenAI models, because when it is, with OpenAI models, because when it is, this model is going to be crazy on it, this model is going to be crazy on it, this model is going to be crazy on it, because it can already figure out how to because it can already figure out how to because it can already figure out how to navigate computer use totally fine. If navigate computer use totally fine. If navigate computer use totally fine. If it can suddenly do it six times faster, it can suddenly do it six times faster, it can suddenly do it six times faster, it's going to be unbelievably fun. Still it's going to be unbelievably fun. Still it's going to be unbelievably fun. Still not quite as good as Astra, but more not quite as good as Astra, but more not quite as good as Astra, but more than good enough that for the price than good enough that for the price than good enough that for the price difference, I wouldn't even think twice difference, I wouldn't even think twice difference, I wouldn't even think twice about it. But, as I mentioned before, about it. But, as I mentioned before, about it. But, as I mentioned before, there are certain things I would still there are certain things I would still there are certain things I would still occasionally use Astra for that I am occasionally use Astra for that I am occasionally use Astra for that I am more than happy to use Soul for. One of more than happy to use Soul for. One of more than happy to use Soul for. One of those things is deep code reviews. I

  10. those things is deep code reviews. I those things is deep code reviews. I have still found OpenAI models and the have still found OpenAI models and the have still found OpenAI models and the like Rottweiler nature where they'll dig like Rottweiler nature where they'll dig like Rottweiler nature where they'll dig into a problem and shake it into tiny into a problem and shake it into tiny into a problem and shake it into tiny pieces until they find every single pieces until they find every single pieces until they find every single thing wrong with it. I find 6.1 Soul to thing wrong with it. I find 6.1 Soul to thing wrong with it. I find 6.1 Soul to be incredibly capable in this particular be incredibly capable in this particular be incredibly capable in this particular way. So, as you can probably guess, I way. So, as you can probably guess, I way. So, as you can probably guess, I had 6.1 Soul do some deep audits on had 6.1 Soul do some deep audits on had 6.1 Soul do some deep audits on Orchestrator V2 and other parts of my Orchestrator V2 and other parts of my Orchestrator V2 and other parts of my real-world codebases. In my Orchestrator real-world codebases. In my Orchestrator real-world codebases. In my Orchestrator V2 audit, it performed nearly V2 audit, it performed nearly V2 audit, it performed nearly identically to Astra. I do believe it identically to Astra. I do believe it identically to Astra. I do believe it was slightly higher a score. Okay, not was slightly higher a score. Okay, not was slightly higher a score. Okay, not in this analysis, but in my other in this analysis, but in my other in this analysis, but in my other analysis, it did actually score very, analysis, it did actually score very, analysis, it did actually score very, very slightly higher. But, it did it at very slightly higher. But, it did it at very slightly higher. But, it did it at about half the price, 297 versus 584. about half the price, 297 versus 584. about half the price, 297 versus 584. Sonnet was still cheaper, and Opus was Sonnet was still cheaper, and Opus was Sonnet was still cheaper, and Opus was slightly cheaper as well. The difference slightly cheaper as well. The difference slightly cheaper as well. The difference being neither of these models were being neither of these models were being neither of these models were anywhere near as thorough with their anywhere near as thorough with their anywhere near as thorough with their analysis. 6.1 Soul dug deep to find analysis. 6.1 Soul dug deep to find analysis. 6.1 Soul dug deep to find things, which is why it was able to get things, which is why it was able to get things, which is why it was able to get a score comparable to Astra, although it a score comparable to Astra, although it a score comparable to Astra, although it did admittedly burn way more tokens. did admittedly burn way more tokens. did admittedly burn way more tokens. Another task I've had a lot of fun Another task I've had a lot of fun Another task I've had a lot of fun testing with is asking the model to find testing with is asking the model to find testing with is asking the model to find opportunities to improve a codebase. In opportunities to improve a codebase. In opportunities to improve a codebase. In this case, to improve D3 code. This is this case, to improve D3 code. This is this case, to improve D3 code. This is the one where Grok 4.7 scored strangely the one where Grok 4.7 scored strangely the one where Grok 4.7 scored strangely well. Of course, Astra scored way better well. Of course, Astra scored way better well. Of course, Astra scored way better at an 83.8 versus the 80.7 from Grok at an 83.8 versus the 80.7 from Grok at an 83.8 versus the 80.7 from Grok 4.7. But, GBD 6.1 Soul hit it out of the 4.7. But, GBD 6.1 Soul hit it out of the 4.7. But, GBD 6.1 Soul hit it out of the park with an 87.4.

  11. park with an 87.4. park with an 87.4. I didn't save all the prices for these I didn't save all the prices for these I didn't save all the prices for these runs. It's been a bit, okay? But, for runs. It's been a bit, okay? But, for runs. It's been a bit, okay? But, for Opus 5.5, it cost five bucks. And for Opus 5.5, it cost five bucks. And for Opus 5.5, it cost five bucks. And for Sonnet 5.5, it cost almost $9. With GBD Sonnet 5.5, it cost almost $9. With GBD Sonnet 5.5, it cost almost $9. With GBD 6.1 Soul, it was $2.15. 6.1 Soul, it was $2.15. 6.1 Soul, it was $2.15. That's the difference. This model's That's the difference. This model's That's the difference. This model's token price is cheaper than Sonnet, but token price is cheaper than Sonnet, but token price is cheaper than Sonnet, but its token utilization is still its token utilization is still its token utilization is still maintaining OpenAI's usual efficiency, maintaining OpenAI's usual efficiency, maintaining OpenAI's usual efficiency, which results in just crazy price to which results in just crazy price to which results in just crazy price to performance. This whole thread was performance. This whole thread was performance. This whole thread was particularly fun because I had Opus 5.5 particularly fun because I had Opus 5.5 particularly fun because I had Opus 5.5 review this model with a different name, review this model with a different name, review this model with a different name, obviously, so it didn't know what it obviously, so it didn't know what it obviously, so it didn't know what it was. I went and edited the history was. I went and edited the history was. I went and edited the history after. And it concluded very quickly after. And it concluded very quickly after. And it concluded very quickly this was a frontier tier model. Its this was a frontier tier model. Its this was a frontier tier model. Its reviews and bug repros matched the fixes reviews and bug repros matched the fixes reviews and bug repros matched the fixes that later merged. Its first draft code that later merged. Its first draft code that later merged. Its first draft code had real bugs, which review bots caught. had real bugs, which review bots caught. had real bugs, which review bots caught. Four reviewers are still running. Local Four reviewers are still running. Local Four reviewers are still running. Local team for code reviews backed frontier team for code reviews backed frontier team for code reviews backed frontier tier again. In a blind 10-model bench on tier again. In a blind 10-model bench on tier again. In a blind 10-model bench on the same prompt, 6.1 Soul placed first the same prompt, 6.1 Soul placed first the same prompt, 6.1 Soul placed first with the 87.4. It found the fish slop with the 87.4. It found the fish slop with the 87.4. It found the fish slop runs and compared those two. It did say runs and compared those two. It did say runs and compared those two. It did say 6.1 souls quality output was slightly 6.1 souls quality output was slightly 6.1 souls quality output was slightly below Astra's as well as the two open AI below Astra's as well as the two open AI below Astra's as well as the two open AI models with Opus 5.5 and sauna 5.5, models with Opus 5.5 and sauna 5.5, models with Opus 5.5 and sauna 5.5, which we will absolutely show you in a which we will absolutely show you in a which we will absolutely show you in a bit. But I do want to call out the price bit. But I do want to call out the price bit. But I do want to call out the price here cuz it only costs $7 to run versus here cuz it only costs $7 to run versus here cuz it only costs $7 to run versus 15 for sauna 5.5 and 50 for Opus. Opus's 15 for sauna 5.5 and 50 for Opus. Opus's 15 for sauna 5.5 and 50 for Opus. Opus's honest tier call was that this model is honest tier call was that this model is honest tier call was that this model is incredible for scoped work, top of the incredible for scoped work, top of the incredible for scoped work, top of the frontier. For find what's wrong and tell frontier. For find what's wrong and tell frontier. For find what's wrong and tell me the truth, I would choose it over me the truth, I would choose it over me the truth, I would choose it over Astra and about level with Opus. For Astra and about level with Opus. For Astra and about level with Opus. For long unattended building, this was below long unattended building, this was below long unattended building, this was below frontier. Follows its process rules even frontier. Follows its process rules even frontier. Follows its process rules even when they stop all progress and it does

  12. when they stop all progress and it does when they stop all progress and it does not ask for help. This I absolutely not ask for help. This I absolutely not ask for help. This I absolutely noticed. I had mentioned before, well, a noticed. I had mentioned before, well, a noticed. I had mentioned before, well, a few times now, that my TS rust port that few times now, that my TS rust port that few times now, that my TS rust port that I'm working with Opus 5.5 is going way I'm working with Opus 5.5 is going way I'm working with Opus 5.5 is going way better than when I was working on that better than when I was working on that better than when I was working on that same port using Astra and soul in the same port using Astra and soul in the same port using Astra and soul in the past. I had that port running for a past. I had that port running for a past. I had that port running for a while with this model and it made no while with this model and it made no while with this model and it made no progress. It burned a shitload of progress. It burned a shitload of progress. It burned a shitload of tokens, but it didn't actually improve tokens, but it didn't actually improve tokens, but it didn't actually improve the compiler at all. Opus was able to the compiler at all. Opus was able to the compiler at all. Opus was able to from scratch restart it and get it from scratch restart it and get it from scratch restart it and get it working in a day after I had spent working in a day after I had spent working in a day after I had spent months and hundreds of thousands of months and hundreds of thousands of months and hundreds of thousands of dollars in tokens with this model as dollars in tokens with this model as dollars in tokens with this model as well as with Astra and 5.6 soul. Opus well as with Astra and 5.6 soul. Opus well as with Astra and 5.6 soul. Opus did it in like a grand in like a night did it in like a grand in like a night did it in like a grand in like a night with just two subscriptions with the with just two subscriptions with the with just two subscriptions with the Claude plan. So, for unattended long Claude plan. So, for unattended long Claude plan. So, for unattended long like heavy rewrite type stuff, Anthropic like heavy rewrite type stuff, Anthropic like heavy rewrite type stuff, Anthropic is just comically far ahead right now. is just comically far ahead right now. is just comically far ahead right now. And it also didn't have great judgment And it also didn't have great judgment And it also didn't have great judgment when I was using it for managing my when I was using it for managing my when I was using it for managing my fleet. And for those wondering, my fleet fleet. And for those wondering, my fleet fleet. And for those wondering, my fleet is all the computers I use for running is all the computers I use for running is all the computers I use for running all my agents in code cuz one computer all my agents in code cuz one computer all my agents in code cuz one computer is far from enough. I don't run any of is far from enough. I don't run any of is far from enough. I don't run any of them on this MacBook now. So, when I use them on this MacBook now. So, when I use them on this MacBook now. So, when I use this model to manage the fleet, it made this model to manage the fleet, it made this model to manage the fleet, it made a couple dumb mistakes here and there. a couple dumb mistakes here and there. a couple dumb mistakes here and there. To be fair, so is Opus. Astra's the only To be fair, so is Opus. Astra's the only To be fair, so is Opus. Astra's the only one that hasn't really made too many of one that hasn't really made too many of one that hasn't really made too many of those dumb mistakes, but like I'm going those dumb mistakes, but like I'm going those dumb mistakes, but like I'm going to be so real, I am entirely done using to be so real, I am entirely done using to be so real, I am entirely done using Astra after this model. After I had Opus Astra after this model. After I had Opus Astra after this model. After I had Opus do all of this review, I asked it how do all of this review, I asked it how do all of this review, I asked it how much does it think this model should much does it think this model should much does it think this model should cost. It guessed $5 per mil in, 50 cents cost. It guessed $5 per mil in, 50 cents cost. It guessed $5 per mil in, 50 cents cashed, and 30 per mil out, putting it cashed, and 30 per mil out, putting it cashed, and 30 per mil out, putting it at Opus's prices roughly. It said that at Opus's prices roughly. It said that at Opus's prices roughly. It said that cuz it's performing like Opus, its speed cuz it's performing like Opus, its speed cuz it's performing like Opus, its speed should add a premium cuz it is quite should add a premium cuz it is quite should add a premium cuz it is quite fast. And it's not a pro model, which is fast. And it's not a pro model, which is fast. And it's not a pro model, which is where expects those higher like $100 out where expects those higher like $100 out where expects those higher like $100 out tiered pricing things to come. Pro model

  13. tiered pricing things to come. Pro model tiered pricing things to come. Pro model isn't really a thing anymore. We just isn't really a thing anymore. We just isn't really a thing anymore. We just use Fable and Aster, but you get the use Fable and Aster, but you get the use Fable and Aster, but you get the idea. This is the funniest part of the idea. This is the funniest part of the idea. This is the funniest part of the whole thread though. If OpenAI wants whole thread though. If OpenAI wants whole thread though. If OpenAI wants people to adopt it, I would expect $3 people to adopt it, I would expect $3 people to adopt it, I would expect $3 per mil in, 30 cents for cashed, and $20 per mil in, 30 cents for cashed, and $20 per mil in, 30 cents for cashed, and $20 instead. That would still be a fair instead. That would still be a fair instead. That would still be a fair price for what it does. To which I price for what it does. To which I price for what it does. To which I responded, if I told you it was $2 and responded, if I told you it was $2 and responded, if I told you it was $2 and $10 out and 10 cents per mil cash read, $10 out and 10 cents per mil cash read, $10 out and 10 cents per mil cash read, what would you think? I'd call that very what would you think? I'd call that very what would you think? I'd call that very aggressive pricing. For how you use it, aggressive pricing. For how you use it, aggressive pricing. For how you use it, it costs about a quarter of what I it costs about a quarter of what I it costs about a quarter of what I guessed. The cash price does most of the guessed. The cash price does most of the guessed. The cash price does most of the work. Agent workloads are 96% cash work. Agent workloads are 96% cash work. Agent workloads are 96% cash reads, so 10 cents for cash reads reads, so 10 cents for cash reads reads, so 10 cents for cash reads matters more than the $2 and $10 matters more than the $2 and $10 matters more than the $2 and $10 headline prices. headline prices. headline prices. When it looked at all of my sessions, When it looked at all of my sessions, When it looked at all of my sessions, its price guess would have been $5,700, its price guess would have been $5,700, its price guess would have been $5,700, but after looking at these new prices, but after looking at these new prices, but after looking at these new prices, it redid the math and it would have been it redid the math and it would have been it redid the math and it would have been $1,550. $1,550. $1,550. That is a massive decrease. And for all That is a massive decrease. And for all That is a massive decrease. And for all my PR review type tasks, it was my PR review type tasks, it was my PR review type tasks, it was expecting those to be up to $10 and it's expecting those to be up to $10 and it's expecting those to be up to $10 and it's actually only up to $3. And that's for actually only up to $3. And that's for actually only up to $3. And that's for like heavy PRs with tens of thousands of like heavy PRs with tens of thousands of like heavy PRs with tens of thousands of lines of code. According to Opus, so lines of code. According to Opus, so lines of code. According to Opus, so don't blame me, blame Opus for saying don't blame me, blame Opus for saying don't blame me, blame Opus for saying this. First off, Opus says it becomes this. First off, Opus says it becomes this. First off, Opus says it becomes the default model for scoped work.

  14. the default model for scoped work. the default model for scoped work. Second off, it says that bloated system Second off, it says that bloated system Second off, it says that bloated system prompts barely matter anymore because of prompts barely matter anymore because of prompts barely matter anymore because of the cash pricing. It's just noise. the cash pricing. It's just noise. the cash pricing. It's just noise. Third, it says long loops are still a Third, it says long loops are still a Third, it says long loops are still a bad idea, but not because of money. It's bad idea, but not because of money. It's bad idea, but not because of money. It's cuz according to it, the TSR port wasted cuz according to it, the TSR port wasted cuz according to it, the TSR port wasted four days and made no progress at all. four days and made no progress at all. four days and made no progress at all. And Opus even said they'd be suspicious And Opus even said they'd be suspicious And Opus even said they'd be suspicious of it lasting. They expect this price to of it lasting. They expect this price to of it lasting. They expect this price to go up in the future. I cannot fathom go up in the future. I cannot fathom go up in the future. I cannot fathom OpenAI ever increasing the price for a OpenAI ever increasing the price for a OpenAI ever increasing the price for a model, but Opus thinking they will is model, but Opus thinking they will is model, but Opus thinking they will is hilarious and shows just how good a hilarious and shows just how good a hilarious and shows just how good a value the model is. I love this call out value the model is. I love this call out value the model is. I love this call out here. GPT-4 1 sold did three rounds of here. GPT-4 1 sold did three rounds of here. GPT-4 1 sold did three rounds of work for about half the cost of Sonnet's work for about half the cost of Sonnet's work for about half the cost of Sonnet's single round. A lot of this comes down single round. A lot of this comes down single round. A lot of this comes down to how context was managed. Well, to how context was managed. Well, to how context was managed. Well, because 6.1 sold is much more efficient, because 6.1 sold is much more efficient, because 6.1 sold is much more efficient, so it's not doing as many calls that so it's not doing as many calls that so it's not doing as many calls that bloat the context. It's not outputting bloat the context. It's not outputting bloat the context. It's not outputting as many tokens that are like building up as many tokens that are like building up as many tokens that are like building up over time. So, the average number of over time. So, the average number of over time. So, the average number of tokens being read per request was only tokens being read per request was only tokens being read per request was only around 110,000 tokens versus 360,000 for around 110,000 tokens versus 360,000 for around 110,000 tokens versus 360,000 for Sonnet 5. The result was that Soul used Sonnet 5. The result was that Soul used Sonnet 5. The result was that Soul used under half as many input tokens as under half as many input tokens as under half as many input tokens as Sonnet, making this model significantly Sonnet, making this model significantly Sonnet, making this model significantly more efficient. Speaking of efficiency, more efficient. Speaking of efficiency, more efficient. Speaking of efficiency, I want to talk about these deep SW I want to talk about these deep SW I want to talk about these deep SW scores a tiny bit more because this is a scores a tiny bit more because this is a scores a tiny bit more because this is a weird bench for me to a fork and include weird bench for me to a fork and include weird bench for me to a fork and include in these things. I actually did it for a in these things. I actually did it for a in these things. I actually did it for a different reason, not to compare against different reason, not to compare against different reason, not to compare against 6.1 Soul, but to compare against a new 6.1 Soul, but to compare against a new 6.1 Soul, but to compare against a new release from OpenRouter, Jev Router.

  15. release from OpenRouter, Jev Router. release from OpenRouter, Jev Router. OpenRouter added Jev Router to try and OpenRouter added Jev Router to try and OpenRouter added Jev Router to try and optimize costs with your requests, and I optimize costs with your requests, and I optimize costs with your requests, and I thought it was an incredibly stupid thought it was an incredibly stupid thought it was an incredibly stupid idea. Once I started running it against idea. Once I started running it against idea. Once I started running it against benchmarks, I confirmed it's an benchmarks, I confirmed it's an benchmarks, I confirmed it's an incredibly stupid idea. It turns out a incredibly stupid idea. It turns out a incredibly stupid idea. It turns out a model that cannot reason, that is given model that cannot reason, that is given model that cannot reason, that is given a prompt into no context, cannot make a a prompt into no context, cannot make a a prompt into no context, cannot make a good decision around how hard the good decision around how hard the good decision around how hard the problem is. And Jev Router ended up problem is. And Jev Router ended up problem is. And Jev Router ended up being DeepSeek V4.1 Flash Router for the being DeepSeek V4.1 Flash Router for the being DeepSeek V4.1 Flash Router for the vast majority of its runs. Around 60% of vast majority of its runs. Around 60% of vast majority of its runs. Around 60% of all the requests went straight to all the requests went straight to all the requests went straight to DeepSeek 4.1 Flash. So, didn't like it DeepSeek 4.1 Flash. So, didn't like it DeepSeek 4.1 Flash. So, didn't like it that much. It also routes to other that much. It also routes to other that much. It also routes to other smarter models, which should give it smarter models, which should give it smarter models, which should give it more of an advantage, but it ended up more of an advantage, but it ended up more of an advantage, but it ended up being more expensive than GPT-6 Astro being more expensive than GPT-6 Astro being more expensive than GPT-6 Astro was on low, while also taking four to was on low, while also taking four to was on low, while also taking four to five times longer cuz 6 Astro low took five times longer cuz 6 Astro low took five times longer cuz 6 Astro low took 4.6 minutes, and Jev Router took 20. Jev 4.6 minutes, and Jev Router took 20. Jev 4.6 minutes, and Jev Router took 20. Jev Router's average task took 104 steps, Router's average task took 104 steps, Router's average task took 104 steps, whereas GPT-6 Astro took 19. You get the whereas GPT-6 Astro took 19. You get the whereas GPT-6 Astro took 19. You get the idea, it wasn't very good. But, the idea, it wasn't very good. But, the idea, it wasn't very good. But, the whole point of Jev Router is that it whole point of Jev Router is that it whole point of Jev Router is that it would be as cheap as possible to get a would be as cheap as possible to get a would be as cheap as possible to get a certain score. That was the promise on certain score. That was the promise on certain score. That was the promise on the tin. Whether or not you believe them the tin. Whether or not you believe them the tin. Whether or not you believe them is up to you, not me. I think it's is up to you, not me. I think it's is up to you, not me. I think it's Regardless, Jev Router was Regardless, Jev Router was Regardless, Jev Router was routing to DeepSeek 4.1 Flash for the routing to DeepSeek 4.1 Flash for the routing to DeepSeek 4.1 Flash for the majority of its requests. Despite Jev majority of its requests. Despite Jev majority of its requests. Despite Jev Router routing to the cheapest possible Router routing to the cheapest possible Router routing to the cheapest possible small open-weight models from whatever small open-weight models from whatever small open-weight models from whatever provider will give it away for free. 6.1 provider will give it away for free. 6.1 provider will give it away for free. 6.1 Sonnet low got the same score for an Sonnet low got the same score for an Sonnet low got the same score for an eighth the price. OpenAI is here to eighth the price. OpenAI is here to eighth the price. OpenAI is here to destroy any wins anyone else has in destroy any wins anyone else has in destroy any wins anyone else has in terms of efficiency. Completing this terms of efficiency. Completing this terms of efficiency. Completing this bench in 4.8 minutes for 21 cents with

  16. bench in 4.8 minutes for 21 cents with bench in 4.8 minutes for 21 cents with the second highest score I've ever seen the second highest score I've ever seen the second highest score I've ever seen on it is a massive achievement, tying on it is a massive achievement, tying on it is a massive achievement, tying Opus 5.5, which took 50 minutes per task Opus 5.5, which took 50 minutes per task Opus 5.5, which took 50 minutes per task on max. A tenth the time and a on max. A tenth the time and a on max. A tenth the time and a seventieth the price for the same score. seventieth the price for the same score. seventieth the price for the same score. If your work fits within the things 6.1 If your work fits within the things 6.1 If your work fits within the things 6.1 Sonnet does well, you should probably Sonnet does well, you should probably Sonnet does well, you should probably use it for everything. But, if your work use it for everything. But, if your work use it for everything. But, if your work doesn't fit in it particularly well, you doesn't fit in it particularly well, you doesn't fit in it particularly well, you should probably keep using Opus and should probably keep using Opus and should probably keep using Opus and maybe give Opus the ability to call 6.1 maybe give Opus the ability to call 6.1 maybe give Opus the ability to call 6.1 Sonnet when it should for various tasks. Sonnet when it should for various tasks. Sonnet when it should for various tasks. I'm almost certainly going to be setting I'm almost certainly going to be setting I'm almost certainly going to be setting things up so that Opus 5.5 can call 6.1 things up so that Opus 5.5 can call 6.1 things up so that Opus 5.5 can call 6.1 Sonnet to do investigation work to try Sonnet to do investigation work to try Sonnet to do investigation work to try and like root cause bugs, to do analysis and like root cause bugs, to do analysis and like root cause bugs, to do analysis of code bases to figure out what things of code bases to figure out what things of code bases to figure out what things need to be touched and why, to help me need to be touched and why, to help me need to be touched and why, to help me triage real-world work, to help me triage real-world work, to help me triage real-world work, to help me review the work that Opus does, and review the work that Opus does, and review the work that Opus does, and more. I'm kind of spoiling the ending more. I'm kind of spoiling the ending more. I'm kind of spoiling the ending here, aren't I? I'm going to keep using here, aren't I? I'm going to keep using here, aren't I? I'm going to keep using Opus 5.5 for now. Before I explain why, Opus 5.5 for now. Before I explain why, Opus 5.5 for now. Before I explain why, let me do the thing that I'm most let me do the thing that I'm most let me do the thing that I'm most excited for. excited for. excited for. Fish slop. Fish slop. Fish slop. The first thing you might have noticed The first thing you might have noticed The first thing you might have noticed is the inclusion of slop in fish slop.

  17. is the inclusion of slop in fish slop. is the inclusion of slop in fish slop. This model did the horrible thing I hate This model did the horrible thing I hate This model did the horrible thing I hate where it surrounded the game in a bunch where it surrounded the game in a bunch where it surrounded the game in a bunch of absolutely garbage UI. And coming to of absolutely garbage UI. And coming to of absolutely garbage UI. And coming to this right after the 5.5 Sonnet demo this right after the 5.5 Sonnet demo this right after the 5.5 Sonnet demo hurts me deeply hurts me deeply hurts me deeply because Sonnet 5.5 did not make graphics because Sonnet 5.5 did not make graphics because Sonnet 5.5 did not make graphics anywhere near this good looking, anywhere near this good looking, anywhere near this good looking, but at least it made a UI that was but at least it made a UI that was but at least it made a UI that was nowhere near this awful. And man, do I nowhere near this awful. And man, do I nowhere near this awful. And man, do I wish the bad UI is where the issues wish the bad UI is where the issues wish the bad UI is where the issues stopped. I will turn on the sound. stopped. I will turn on the sound. stopped. I will turn on the sound. Oh god, it's blaring. It is stunning looking. It is stunning looking. The fish are some of the best. The model The fish are some of the best. The model The fish are some of the best. The model for the sub is way better. The for the sub is way better. The for the sub is way better. The propellers work way better. propellers work way better. propellers work way better. I'm going to mute the sound cuz that is I'm going to mute the sound cuz that is I'm going to mute the sound cuz that is looking pretty bad. I haven't even heard looking pretty bad. I haven't even heard looking pretty bad. I haven't even heard it, honestly. it, honestly. it, honestly. But damn, like But damn, like But damn, like looks beautiful, looks beautiful, looks beautiful, but if you actually are playing it, one but if you actually are playing it, one but if you actually are playing it, one of the first things you'll notice is of the first things you'll notice is of the first things you'll notice is that the movement feels significantly that the movement feels significantly that the movement feels significantly worse than it does in either the Opus or worse than it does in either the Opus or worse than it does in either the Opus or the Sonnet versions that I've demoed in the Sonnet versions that I've demoed in the Sonnet versions that I've demoed in the past.

  18. Yeah, it it moves jank. It also has a Yeah, it it moves jank. It also has a significantly worse frame rate than the significantly worse frame rate than the significantly worse frame rate than the versions from the other models. It does versions from the other models. It does versions from the other models. It does have higher graphic fidelity, so that have higher graphic fidelity, so that have higher graphic fidelity, so that makes sense. Like the models here with makes sense. Like the models here with makes sense. Like the models here with the plants are significantly better than the plants are significantly better than the plants are significantly better than they were with the Sonnet version. The they were with the Sonnet version. The they were with the Sonnet version. The details of the fidelity of the extras in details of the fidelity of the extras in details of the fidelity of the extras in the tank is absolutely hilarious. the tank is absolutely hilarious. the tank is absolutely hilarious. Like Like Like yeah. yeah. yeah. But god damn, I'm so tired of the But god damn, I'm so tired of the But god damn, I'm so tired of the unnecessary text everywhere. This model unnecessary text everywhere. This model unnecessary text everywhere. This model does it worse than almost any I've ever does it worse than almost any I've ever does it worse than almost any I've ever seen before. We got take a breather, seen before. We got take a breather, seen before. We got take a breather, paused, paused, paused, your little world can wait, back to the your little world can wait, back to the your little world can wait, back to the reef, start a new tank, slot 01, feeder reef, start a new tank, slot 01, feeder reef, start a new tank, slot 01, feeder submarine, submarine, submarine, a little underwater chaos, a little underwater chaos, a little underwater chaos, the big little goal, your little the big little goal, your little the big little goal, your little ecosystem, little fish become big ecosystem, little fish become big ecosystem, little fish become big earners, four meals and they're all earners, four meals and they're all earners, four meals and they're all grown up, grown up, grown up, make the family a little bigger, a make the family a little bigger, a make the family a little bigger, a little golden overachiever. It is so little golden overachiever. It is so little golden overachiever. It is so many of these. There's like 20 plus of many of these. There's like 20 plus of many of these. There's like 20 plus of them. them. them. And I promise you guys, as soon as I saw And I promise you guys, as soon as I saw And I promise you guys, as soon as I saw this, I took a screenshot, I sent it to this, I took a screenshot, I sent it to this, I took a screenshot, I sent it to Open AI, and I crashed out in the slack Open AI, and I crashed out in the slack Open AI, and I crashed out in the slack because I cannot fathom how they haven't because I cannot fathom how they haven't because I cannot fathom how they haven't fixed this problem. This model fixed this problem. This model fixed this problem. This model is unacceptably garbage at UI. It has is unacceptably garbage at UI. It has is unacceptably garbage at UI. It has regressed again. And if you're looking regressed again. And if you're looking regressed again. And if you're looking for a model that can make front ends for a model that can make front ends for a model that can make front ends that don't suck, go spend your money that don't suck, go spend your money that don't suck, go spend your money somewhere else because it should not be somewhere else because it should not be somewhere else because it should not be spent here. This model sucks at front spent here. This model sucks at front spent here. This model sucks at front end. It sucks at design. It has no end. It sucks at design. It has no end. It sucks at design. It has no taste. And you're going to have to bring taste. And you're going to have to bring taste. And you're going to have to bring your taste yourself still. But it is your taste yourself still. But it is your taste yourself still. But it is admittedly really good at Blender. If admittedly really good at Blender. If admittedly really good at Blender. If you give it like a screenshot of a thing you give it like a screenshot of a thing you give it like a screenshot of a thing you want it to model in 3D and say, you want it to model in 3D and say, you want it to model in 3D and say, "Hey, you have Blender over the CLI. Go

  19. "Hey, you have Blender over the CLI. Go "Hey, you have Blender over the CLI. Go make this." It will, and it'll do a make this." It will, and it'll do a make this." It will, and it'll do a pretty damn good job. pretty damn good job. pretty damn good job. But I would never have it make the But I would never have it make the But I would never have it make the actual mechanics for my games because it actual mechanics for my games because it actual mechanics for my games because it feels awful to play. It also has like feels awful to play. It also has like feels awful to play. It also has like nowhere near as much gameplay loop. In nowhere near as much gameplay loop. In nowhere near as much gameplay loop. In fact, the first time I tried demoing fact, the first time I tried demoing fact, the first time I tried demoing this before filming, this before filming, this before filming, it just randomly game overed as I was it just randomly game overed as I was it just randomly game overed as I was like getting started in the first 30 like getting started in the first 30 like getting started in the first 30 seconds. It never like said why. seconds. It never like said why. seconds. It never like said why. Actually, I think I technically beat it. Actually, I think I technically beat it. Actually, I think I technically beat it. I also want to like take this version I also want to like take this version I also want to like take this version and hand it to Opus or Sonnet and say, and hand it to Opus or Sonnet and say, and hand it to Opus or Sonnet and say, "Hey, can you make this play better? "Hey, can you make this play better? "Hey, can you make this play better? Because the graphics are good, but the Because the graphics are good, but the Because the graphics are good, but the game sucks." But when you combine how game sucks." But when you combine how game sucks." But when you combine how cheap it was to make this cuz like this cheap it was to make this cuz like this cheap it was to make this cuz like this was $5, I think, to generate. That's was $5, I think, to generate. That's was $5, I think, to generate. That's pretty insane. And if you combine that pretty insane. And if you combine that pretty insane. And if you combine that with like Ultra Fast, if that ever with like Ultra Fast, if that ever with like Ultra Fast, if that ever happens, suddenly you're going to be happens, suddenly you're going to be happens, suddenly you're going to be able to make a game in a few minutes on able to make a game in a few minutes on able to make a game in a few minutes on demand. We're actually now getting to demand. We're actually now getting to demand. We're actually now getting to that threshold where game development is that threshold where game development is that threshold where game development is about to flip upside down because of how about to flip upside down because of how about to flip upside down because of how models are finally understanding models are finally understanding models are finally understanding three-dimensional space and the tooling three-dimensional space and the tooling three-dimensional space and the tooling necessary to do these types of things. necessary to do these types of things. necessary to do these types of things. It's happening. As per usual, I was not It's happening. As per usual, I was not It's happening. As per usual, I was not allowed to put the code I wrote with allowed to put the code I wrote with allowed to put the code I wrote with this model inside of T3 Code or other this model inside of T3 Code or other this model inside of T3 Code or other open-source projects during the testing open-source projects during the testing open-source projects during the testing window. So, I had to use it exclusively window. So, I had to use it exclusively window. So, I had to use it exclusively on my internal projects like Lakebed, as on my internal projects like Lakebed, as on my internal projects like Lakebed, as well as for auditing other work, which well as for auditing other work, which well as for auditing other work, which means I mostly use this for auditing means I mostly use this for auditing means I mostly use this for auditing other work, and I was very impressed.

  20. other work, and I was very impressed. other work, and I was very impressed. This is a real PR I was working on to This is a real PR I was working on to This is a real PR I was working on to fix a bug where my new little work tree fix a bug where my new little work tree fix a bug where my new little work tree like setup window that would appear in a like setup window that would appear in a like setup window that would appear in a new thread in T3 Code would disappear if new thread in T3 Code would disappear if new thread in T3 Code would disappear if you left and came back. I had Claude you left and came back. I had Claude you left and came back. I had Claude Code work on this, but this problem went Code work on this, but this problem went Code work on this, but this problem went pretty deep. So, I wanted to make sure pretty deep. So, I wanted to make sure pretty deep. So, I wanted to make sure that whatever solution I came up with that whatever solution I came up with that whatever solution I came up with was very, very, very well vetted. While was very, very, very well vetted. While was very, very, very well vetted. While I personally still do not trust this I personally still do not trust this I personally still do not trust this model to write the code I'm trying to model to write the code I'm trying to model to write the code I'm trying to land, I absolutely trust it to review land, I absolutely trust it to review land, I absolutely trust it to review things. Ignore the GPT-6 soul there, things. Ignore the GPT-6 soul there, things. Ignore the GPT-6 soul there, just a placeholder. So, when I had 6.1 just a placeholder. So, when I had 6.1 just a placeholder. So, when I had 6.1 Soul look through this, it found real Soul look through this, it found real Soul look through this, it found real problems that were entirely missed by problems that were entirely missed by problems that were entirely missed by Fable and by Opus. First, it called out Fable and by Opus. First, it called out Fable and by Opus. First, it called out that follow-up messages can stay blocked that follow-up messages can stay blocked that follow-up messages can stay blocked after the agent starts, which is very after the agent starts, which is very after the agent starts, which is very annoying if you want to queue a message, annoying if you want to queue a message, annoying if you want to queue a message, and also that recovered setup progress and also that recovered setup progress and also that recovered setup progress was disappearing too early. It figured was disappearing too early. It figured was disappearing too early. It figured all of these things out with a all of these things out with a all of these things out with a combination of reading the code and combination of reading the code and combination of reading the code and analyzing it as well as computer use. It analyzing it as well as computer use. It analyzing it as well as computer use. It was able to prevent me from merging a was able to prevent me from merging a was able to prevent me from merging a real regression in T3 code. So, I real regression in T3 code. So, I real regression in T3 code. So, I literally just copy-pasted those things literally just copy-pasted those things literally just copy-pasted those things to Claude and then told it to take to Claude and then told it to take to Claude and then told it to take another look. Said, "Better, but I'll another look. Said, "Better, but I'll another look. Said, "Better, but I'll still fix two small gaps." Remember, I still fix two small gaps." Remember, I still fix two small gaps." Remember, I can't use this model of the code for can't use this model of the code for can't use this model of the code for this project at the time, so I this project at the time, so I this project at the time, so I copy-pasted that again over to Claude, copy-pasted that again over to Claude, copy-pasted that again over to Claude, and it eventually got it good enough, and it eventually got it good enough, and it eventually got it good enough, and then I finally merged. But, that is and then I finally merged. But, that is and then I finally merged. But, that is what I like this model for, and I cannot what I like this model for, and I cannot what I like this model for, and I cannot wait to push its limits for actually wait to push its limits for actually wait to push its limits for actually coding. Although, I will say from the coding. Although, I will say from the coding. Although, I will say from the code that I did have the misfortune of code that I did have the misfortune of code that I did have the misfortune of reading, it is harder to justify merging reading, it is harder to justify merging reading, it is harder to justify merging this code than it is for code from Opus.

  21. this code than it is for code from Opus. this code than it is for code from Opus. Normally, I would make you guys wait for Normally, I would make you guys wait for Normally, I would make you guys wait for the Opus versus Soul video or the Sonic the Opus versus Soul video or the Sonic the Opus versus Soul video or the Sonic versus Soul video, but I'll just spoil versus Soul video, but I'll just spoil versus Soul video, but I'll just spoil the details now. I like using them in the details now. I like using them in the details now. I like using them in tandem because I find Soul to be way tandem because I find Soul to be way tandem because I find Soul to be way better at reviewing and digging into the better at reviewing and digging into the better at reviewing and digging into the details, but I find Opus a more pleasant details, but I find Opus a more pleasant details, but I find Opus a more pleasant collaborator and significantly better at collaborator and significantly better at collaborator and significantly better at actually implementing code without actually implementing code without actually implementing code without getting blocked constantly throughout getting blocked constantly throughout getting blocked constantly throughout its work. And even now with the Rust its work. And even now with the Rust its work. And even now with the Rust rewrite of TypeScript, I find myself in rewrite of TypeScript, I find myself in rewrite of TypeScript, I find myself in a similar pattern, where I have Opus 5.5 a similar pattern, where I have Opus 5.5 a similar pattern, where I have Opus 5.5 just going and going and going, making just going and going and going, making just going and going and going, making the code base work and what work well, the code base work and what work well, the code base work and what work well, and then I had Soul come in and do an and then I had Soul come in and do an and then I had Soul come in and do an audit. And this is the funniest part. audit. And this is the funniest part. audit. And this is the funniest part. Remember before I said that I had Soul Remember before I said that I had Soul Remember before I said that I had Soul and Astra working on that TS Rust port and Astra working on that TS Rust port and Astra working on that TS Rust port for effectively months? I told Soul to for effectively months? I told Soul to for effectively months? I told Soul to come in, and it got it from 83.7% to come in, and it got it from 83.7% to come in, and it got it from 83.7% to 100% in under a day. I was blown away by 100% in under a day. I was blown away by 100% in under a day. I was blown away by that, that it had somehow unblocked the that, that it had somehow unblocked the that, that it had somehow unblocked the work that Astra and Soul were doing, as work that Astra and Soul were doing, as work that Astra and Soul were doing, as well as 6.1 Soul. And I was absolutely well as 6.1 Soul. And I was absolutely well as 6.1 Soul. And I was absolutely blown away by that, that it had taken blown away by that, that it had taken blown away by that, that it had taken the work that 5.6 Soul, 6.1 Soul, and the work that 5.6 Soul, 6.1 Soul, and the work that 5.6 Soul, 6.1 Soul, and Astra had done over months, and got it Astra had done over months, and got it Astra had done over months, and got it unblocked where it had been stuck for unblocked where it had been stuck for unblocked where it had been stuck for weeks, and finished it. I was much more weeks, and finished it. I was much more weeks, and finished it. I was much more blown away when I had 61 Soul take a blown away when I had 61 Soul take a blown away when I had 61 Soul take a look at that work and critique it. And look at that work and critique it. And look at that work and critique it. And what it brought up was that of the 1.8 what it brought up was that of the 1.8 what it brought up was that of the 1.8 million lines of code, 1.3 million were million lines of code, 1.3 million were million lines of code, 1.3 million were not being used. The reason was because not being used. The reason was because not being used. The reason was because Opus concluded whole of the code from Opus concluded whole of the code from Opus concluded whole of the code from all the other agents was useless slop all the other agents was useless slop all the other agents was useless slop that had no chance of being recovered, that had no chance of being recovered, that had no chance of being recovered, and it chose to rewrite it from scratch and it chose to rewrite it from scratch and it chose to rewrite it from scratch itself in another crate. So, on one itself in another crate. So, on one itself in another crate. So, on one hand, the only reason the code worked hand, the only reason the code worked hand, the only reason the code worked was Opus, but on the other hand, the was Opus, but on the other hand, the was Opus, but on the other hand, the only reason the slop was still around

  22. only reason the slop was still around only reason the slop was still around was also Opus. So, I had to have this was also Opus. So, I had to have this was also Opus. So, I had to have this model come in and clean up the mess that model come in and clean up the mess that model come in and clean up the mess that other OpenAI models had made because other OpenAI models had made because other OpenAI models had made because Opus didn't even notice the mess was Opus didn't even notice the mess was Opus didn't even notice the mess was still there. What I'm trying to say is still there. What I'm trying to say is still there. What I'm trying to say is this model absolutely has a place in this model absolutely has a place in this model absolutely has a place in your workflows. It could probably even your workflows. It could probably even your workflows. It could probably even be your default coding model, and you be your default coding model, and you be your default coding model, and you wouldn't have too many issues with it, wouldn't have too many issues with it, wouldn't have too many issues with it, but I still find Opus to be a better but I still find Opus to be a better but I still find Opus to be a better collaborator overall. That said, I have collaborator overall. That said, I have collaborator overall. That said, I have almost no reason to use Sonnet anymore almost no reason to use Sonnet anymore almost no reason to use Sonnet anymore because this will effectively take its because this will effectively take its because this will effectively take its place. And you bet your butt the moment place. And you bet your butt the moment place. And you bet your butt the moment this model comes out, I'll be going and this model comes out, I'll be going and this model comes out, I'll be going and making adjustments inside of my Claude making adjustments inside of my Claude making adjustments inside of my Claude config because I already have it set up config because I already have it set up config because I already have it set up so that I can use Soul inside of Claude so that I can use Soul inside of Claude so that I can use Soul inside of Claude code because I want to make sure Opus code because I want to make sure Opus code because I want to make sure Opus knows this is the model to have review knows this is the model to have review knows this is the model to have review its work and investigate the things its work and investigate the things its work and investigate the things going on in the code base. This is a going on in the code base. This is a going on in the code base. This is a damn good model, and I'm really happy to damn good model, and I'm really happy to damn good model, and I'm really happy to have it. I wish we had something bigger, have it. I wish we had something bigger, have it. I wish we had something bigger, smarter, and more capable overall. I was smarter, and more capable overall. I was smarter, and more capable overall. I was really hoping for something to truly really hoping for something to truly really hoping for something to truly dethrone Opus 5.5 as my daily driver. dethrone Opus 5.5 as my daily driver. dethrone Opus 5.5 as my daily driver. This isn't it, and I'm not planning on This isn't it, and I'm not planning on This isn't it, and I'm not planning on canceling any of my Claude subs as a canceling any of my Claude subs as a canceling any of my Claude subs as a result of this release, but I am result of this release, but I am result of this release, but I am planning on taking a lot more of my planning on taking a lot more of my planning on taking a lot more of my Codex subs in my day-to-day work, Codex subs in my day-to-day work, Codex subs in my day-to-day work, admittedly in Claude code. This is an admittedly in Claude code. This is an admittedly in Claude code. This is an awesome release and an unbelievable awesome release and an unbelievable awesome release and an unbelievable price for what you're getting, but this price for what you're getting, but this price for what you're getting, but this does potentially mark the start of the does potentially mark the start of the does potentially mark the start of the end of the subsidization era. So, make end of the subsidization era. So, make end of the subsidization era. So, make sure you're subscribed so that you can sure you're subscribed so that you can sure you're subscribed so that you can be here when I cover all of that and be here when I cover all of that and be here when I cover all of that and more. God, I hope this isn't get me more. God, I hope this isn't get me more. God, I hope this isn't get me canceled online. I have no idea how canceled online. I have no idea how canceled online. I have no idea how others feel about this beyond like a others feel about this beyond like a others feel about this beyond like a handful of early access testers I've handful of early access testers I've handful of early access testers I've talked to. I legitimately don't know if talked to. I legitimately don't know if talked to. I legitimately don't know if you all are going to love it or hate it you all are going to love it or hate it you all are going to love it or hate it or land somewhere between. I will know or land somewhere between. I will know or land somewhere between. I will know in a few hours, I guess, cuz it's a in a few hours, I guess, cuz it's a in a few hours, I guess, cuz it's a Yeah, it's 3:00 in the morning. I am Yeah, it's 3:00 in the morning. I am Yeah, it's 3:00 in the morning. I am going to go to bed now.

  23. going to go to bed now. going to go to bed now. Uh, Uh, Uh, this is a tiring one. Hopefully I did a this is a tiring one. Hopefully I did a this is a tiring one. Hopefully I did a good job. Let me know in the comments. good job. Let me know in the comments. good job. Let me know in the comments. And until next time, And until next time, And until next time, peace nerds. peace nerds. peace nerds. God, I'm so dead. God, I'm so dead. God, I'm so dead. Ah.

No summary available yet.

View original episode ↗