Latest Posts (20 found)
Unsung Today

“Microsoft’s most ambitious attempt to reinvent the start screen”

From Marton Barcza at TechAltar, a good 12-minute video analyzing what went wrong with the famed Live Tiles that Microsoft was pushing throughout most of the 2010s: = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/microsofts-most-ambitious-attempt-to-reinvent-the-start-screen/yt1-play.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/microsofts-most-ambitious-attempt-to-reinvent-the-start-screen/yt1-play.1600w.avif" type="image/avif"> Now, among Windows Phone fans the pervasive opinion is that Live Tiles have failed because Microsoft did a poor job with them, which I think is at least partially true – after all, even many key Microsoft apps like Skype had broken Live Tiles half the time, the Windows 10 start menu was filled by Microsoft with Live Tiles that were clearly just animated ads rather than showing you something actually useful, and the company of course never built out any of their advanced interactive concepts either. But while all of that might be true, the fact that every major company has walked away from this idea and nobody else has picked it up since, means that there probably are more fundamental problems with this idea. And I can think of three distinct ones: Barcza goes on to talk about some specific interactions and problems, and arrives at the conclusion that Live Tiles ended up at this unpleasant intersection where they’re tried to be icons, widgets, and notifications all at once, not doing a particularly great job at either task. That, and there were also challenges with the design primarily being mobile-first, and struggling surviving a jump to a large screen. (Live Tiles were abandoned in the 2017, at which point desktop Windows reverted back to what it was before, and the mobile/​tablet lines were altogether disbanded.) What I found interesting in watching this today is that it feels Apple has made some similar mistakes in their Liquid Glass approach and wider platform unification desires (see: macOS Settings). Both these and Live Tiles attempted to create A System Of Systems, and both ended up occasionally feeling like they awkwardly crowbarred some wider concepts into places where they didn’t truly belong. If it helps with your further research, I understand Live Tiles were part of a bigger UI effort called Metro and started in earnest with Windows Phone . #apple #change management #complexity #touch #windows #youtube user experience, form factors, and horizontal integration.

0 views

Outburst

Just had a little trip up to my favorite city of Juneau, Alaska. So lovely. Here’s just one story from the trip. There is a Big Thing going on there right now which is essentially an impending flood that is going to happen any day now. It’s an “Outburst Flood” or GLOF (“glacial lake outburst flood”). The weather.gov page explains in one graphic: So there is the famous Mendenhall glacier. Mendenhall lake, formed by it’s melt, is like 30 minutes from downtown Juneau, so it’s quite accessible for even cruise ship visitors to go get a look. Suicide Basin fills up with water during the warm months from glacial melt, and the Mendenhall glacier acts as an ice wall between it and Mendenhall lake. But at some point, the basin gets so full, the water from it starts flowing underneath/through the glacier into Mendenhall lake. There is a river from Mendenhall lake out to the ocean. When Mendenhall lake starts rising from the basin runoff, it goes out that river, and that river rages for a good couple of days. Hence the flooding. That river? Half the people in Juneau live in the Mendenhall Valley where that river runs directly through. Now it’s not all under evacuation flood watch, but some of it certainly is. And it’s rather unknown how bad any given year is going to be. It’s been very bad: My friend Justin lives in the Mendenhall Valley, not far from the river, and when it goes, he can loudly hear the river raging. As I type, we’re just days away from the prediction of August 8-12 from the basin draining. Justin and I climbed up Thunder Mountain (oof, it was actually 8.5 miles and 3,800 ft) and from the summit had a great look down at the entire valley, seeing the whole river area at once. Driving around there on roads by the river, you can see how the city has put up huge wall embankment things to hopefully stop the worst of the damage: Seems like it’s fairly unknown if it’s really going to work. If you’re interested, KTOO did a mini podcast on it last year getting into some details. Like one of the proposed solutions are literally bombing the glacier 🤔. Apparently now it’s leaning more toward the Army Corps of Engineers digging a drainage hole of sorts, but I don’t think any of it is totally sorted out yet.

0 views

Canadian Man Pleads Guilty in Snowflake Extortions

A 26-year-old Canadian man once described as one of the most consequential cybercrime threat actors of 2024 has pleaded guilty to computer fraud and conspiracy to hack and extort more than 165 organizations that used the cloud provider Snowflake . Connor Riley Moucka , of Kitchener, Ontario, also admitted to stealing call and text history records of more than 100 million AT&T customers. A surveillance photo of Connor Riley Moucka, a.k.a. “Judische” and “Waifu,” dated Oct 21, 2024, 9 days before Moucka’s arrest. This image was included in an affidavit filed by an investigator with the Royal Canadian Mounted Police (RCMP). The U.S. Justice Department said between February and October 2024, Moucka and co-conspirators used stolen login credentials to steal cloud-hosted data belonging to at least 165 customers of a U.S.-based software-as-a-service company. The hackers targeted stolen credentials for Snowflake customer accounts that did not enforce multi-factor authentication, and extorted or attempted to extort a host of well-known companies, including TicketMaster, Lending Tree, Advance Auto Parts and Neiman Marcus. Snowflake responded to the data thefts by increasing password complexity requirements and enforcing multi-factor authentication. Moucka adopted new nicknames frequently — sometimes operating multiple identities concurrently — but two of his best-known monikers were “ Judische ” and “ Waifu .” Judische’s admitted role in the Snowflake data thefts was first documented by KrebsOnSecurity in a September 2024 story about the overlap between Western, English-speaking cybercriminals and extremist groups that harass and extort minors into harming themselves or others. That September 2024 story identified Judische as a software engineer from Ontario who has been involved in numerous data breaches and voice phishing attacks against U.S. companies since at least 2020. A little more than a month later, Canadian authorities arrested Moucka on a provisional warrant from the United States. The government says Moucka and others used their unauthorized access to steal billions of sensitive customer records and download terabytes of information, “including individuals’ non-content call and text history records, banking and other financial information, payroll records, Drug Enforcement Administration (DEA) registration numbers, driver’s license numbers, passport numbers, social security numbers and other personally identifiable information. They then extorted victims by threatening to publish data online.” Moucka also threatened and harassed government officials and security researchers who were helping to track him down. The Justice Department said the conspirators made over $2.5 million in ransom payments, and that in at least one instance, Moucka re-extorted a victim with threats of further disclosure of the victim’s stolen data. “Moucka used the stolen data of a government officer and members of a then-former government officer’s immediate family in this re-extortion attempt,” reads a statement from the Justice Department. One of Moucka’s admitted co-conspirators is Cameron “Kiberphant0m” Wagenius , a U.S. Army soldier who pleaded guilty in July 2025 to extorting AT&T and Verizon for their customer account data. Less than a month before Wagenius’s arrest, KrebsOnSecurity published  a deep dive  into Kiberphant0m’s various Telegram and Discord identities over the years, revealing how the owner of the accounts told others they were in the Army and stationed in South Korea. One of several selfies on the Facebook page of Cameron Wagenius. Kiberphant0m also re-extorted victims. Immediately following Moucka’s arrest, Kiberphant0m posted on hacker forums what he claimed were the AT&T call logs for then President-elect Donald Trump and for then Vice President Kamala Harris, as well schematics allegedly stolen from the U.S. National Security Agency (NSA). Wagenius is set to be sentenced on September 3, 2026. The government says he faces a maximum penalty of 20 years in prison for conspiracy to commit wire fraud, a maximum penalty of five years in prison for extortion in relation to computer fraud, and a mandatory two-year sentence consecutive to any other prison time for aggravated identity theft. The third alleged co-conspirator is John Erin Binns , 26, an elusive American man who fled the United States after being indicted for his admitted role in a 2021 breach at T-Mobile that exposed the personal information of at least 76 million customers. Sources close to the investigation said Binns, also known as “ IRDev ” and “ IntelSecrets ,” was until recently incarcerated in a Turkish prison, but that he has since been released and has resurfaced online. Those sources said Binns also recently obtained Turkish citizenship, and under Turkish law a citizen cannot be extradited to a foreign country. An image of a passport that Binns shared in an email to KrebsOnSecurity in Feb. 2023. Moucka pleaded guilty to four criminal counts, including computer fraud, wire fraud, aggravated identity theft, and conspiracy. He is slated to be sentenced on Oct. 27 and faces a mandatory minimum penalty of two years in prison on the aggravated identity theft count, as well as a maximum penalty of 30 years in prison on the remaining counts. Ultimately, it will be up the federal judge how much time Moucka actually serves for his extensive cybercriminal rap sheet. For an interview with Moucka prior to his arrest and a deeper look at Binns, see our original report on Moucka’s arrest .

0 views

seeing the person behind the piece

Edit: This post was being rapidly upvoted by bots for some reason. The issue has been fixed and the fake upvotes removed. Something I noticed as a consequence of online culture and platform design, especially on social media: It is very common online to follow an account not for the person behind it, but the way they present and discuss a topic or niche, even if they aren’t doing this with a professional intent or style. Then when the topics shift or we get tired of the person, people try to find someone else that scratches the itch. It’s like we use people online to engage with topics by proxy, and they are rather replaceable. We don’t just watch crochet videos, we watch this specific person crochet, but not for her as a person, and so on. People seem to have roles to fill in terms of who they follow, and they’ll move on if the role is no longer adequately filled, or the roles change. Feeds are highly curated; too many posts considered off-topic and you are ruining the vibe. In my experience, on the personal web, I really enjoy the culture of people being genuinely interested in learning about someone else. Seems like the general question is “ What’s that person up to? ” There is a wish to see life from the perspective of them, read their diary, explore their interests through their lens, and an acceptance that this person doesn’t necessarily write for you, but more likely for themselves (or at least, also for themselves). From the responses I get via e-mail or blog posts, people seem to respect that this is the specific situation of a stranger, and they’ll compare it with their own. There is a distinct separation. Yet, when my posts breach containment and land on link aggregators or social media, a different crowd shows up. In my experience, people coming from, and heavily using, social media that lets you reshare (reblog/retweet etc) a post onto your own profile tend to read posts as if they are meant to be self-inserts. What I mean is: Every post or article they come across is read as if they are judging whether this crosses a threshold of when they feel like sharing it onto their profiles to be like “ that’s soooo me! ”. Because that’s what happens with, I’d say, most reshares you put your name and picture to: You adopt what this person said and say it too. People find your profile and see everything you reshared as an endorsement, as something you re-say with your own voice. And when this type of person reads my posts, it feels like they do so with tunnel vision, with completely disregarding anything that doesn’t fit into how they see themselves and their own situation. They just wanna plunder my posts for parts that fit enough, so they can plaster it across their feeds as something that represents them, and are mad when there is something in there that is too different from their experience, something they wouldn’t co-sign… and therefore just ignore it or dismiss it, or accuse me of lying. Their responses, while almost never super rude, are at least not empathetic enough in the sense that they write criticisms as if I told a lie about their own life; a life that is extremely different to mine. It seems like for this type of person, the human being behind it doesn’t count at all. Every “content” someone puts out is a meme; further ammo to use to further their view on their own situation, or as self-promotion on their own internet presences. The source is forgettable and irrelevant, because in 5 minutes, they’ll find the next thing to consume. Single use entertainment. It’s just words written by someone else to take into their mouth and spit at the feed 1 , and maybe get upvotes and likes for something they had no hand in. I wouldn’t even mind as much, if it wouldn’t result in a completely warped picture of what I wrote about or what the conclusion is, while posing self-confidently as an expert in the comments trying to give advice or “correct” my experience. And people challenging those is rare because everyone is tired of beefing with random strangers over semantics. Did all this come from the way we are expected to engage with professional content creators who decidedly make mass-appeal content you are support to insert yourself in, and we end up applying it to almost everyone online? How do we feel being treated like a TV channel or magazine? You can be an inspiration, a pastime and distraction to someone; how does that make you feel? Can you withstand the pressure to box yourself in so you are more fitting for the role? Are you beating yourself up for falling short of a label you applied to your blog? Are you already thinking of the brand you have inadvertently built? Are you ready to tear it down over and over again? Published 06 Aug, 2026 Feeds I cannot even properly access or see, because they are all on extremely locked down Mastodon instances which are either completely inaccessible or unusable (bad or no search/content navigation) without an account, which genuinely pisses me off a lot. It's similar to how much info is locked inside Discord servers. You all profit from posting easily accessible public stuff like my blog, but the people create that stuff out cannot even take a peek what you're saying about them. So much for the "better alternatives" and being "open". It's not more open to me than Instagram or X, who also wall their content to death. At least be so kind and send me a direct link via mail, since search engines don't even pick it up either. ↩ Feeds I cannot even properly access or see, because they are all on extremely locked down Mastodon instances which are either completely inaccessible or unusable (bad or no search/content navigation) without an account, which genuinely pisses me off a lot. It's similar to how much info is locked inside Discord servers. You all profit from posting easily accessible public stuff like my blog, but the people create that stuff out cannot even take a peek what you're saying about them. So much for the "better alternatives" and being "open". It's not more open to me than Instagram or X, who also wall their content to death. At least be so kind and send me a direct link via mail, since search engines don't even pick it up either. ↩

0 views

Super Mario Derivations

One of the most surprising aspects of the Nix language is that it is lazy , especially if you have never used a lazy language before. This laziness is what makes much of Nixpkgs possible, and its complexity. One of the simplest ways to observe the laziness is by understanding that only the attributes you access are evaluated. The more whackier version of this is you can have endless recursion in an attribute set. Nixpkgs is filled with these bottomless attribute sets: The same store path every time. contains itself, and so does every package set inside it. 🤯 If laziness is what lets a recursive attribute set terminate, then the recursion doesn’t have to bottom out at all : That attribute set is infinitely deep. Indexing three levels into it costs exactly three levels of evaluation, and the rest of the infinite tree is never built because nobody asked. So an attribute path is a walk through a lazily-generated tree. Which made me wonder: what if the attribute path were input to something ? 🤔 I decided to take that idea and make the attribute path a sequence of button presses in Super Mario Bros. 3 . Each node in the tree is a frame of the game, and each child is a button press that produces a new frame. Game states are recursive by nature. is right + B, which in Super Mario Bros. 3 is “run right”. is run and jump. The output is the frame you’d be looking at if you’d pressed those buttons in that order, on real hardware, in that game. 1 Append anywhere along the path and you get the whole run stitched into a recording: The coolest thing though is that every one of those frames is a separate derivation in my store . The code is at fzakaria/nes-nix . It is generalized and the ROM is a flake input you point wherever you like for any other game. The flake computes a derivation based on the attribute path such that each press is its own derivation, and it takes the previous press’s savestate as an input . Each derivation never re-emulates its ancestors’ frames. 2 The practical consequence is that the store becomes the emulator’s savestate history: Branching off the middle of a hundred-press run costs one press as does appending to the end of it. We can look at it the other way. The dependency graph is the input sequence, so we can ask Nix what buttons produced a frame: So what is actually doing? Almost nothing. Every frame along the path is already sitting in the store as the output of its own press, so the recording never emulates anything. It is a directory of symlinks to the frames for to process. How far can we take this input-sequence game input idea? Nix by default gives out at around 2,400 presses, with: defaults to 10,000 and evaluating each press costs roughly four nested calls. It’s a guard against runaway recursion, not a structural limit, and we can raise it to 10 million and get 20,000 presses: 20,000 presses, takes roughly fourteen seconds to evaluate on my laptop. The cost is linear in the number of presses, and it is roughly 0.7ms “per press”. The next bottleneck though is that the kernel gives out at 21,845 presses on my machine. An attribute path is a single element, and Linux caps the size of the argument list in total and individual arguments. The per-argument limit is 131,072 bytes ( ), and each press is six bytes long ( ), so 21,845 presses is the maximum that can be passed to as a single argument. The escape hatch is to stop passing the run as an argument. and we can feed in the input-sequence as from a file: This produces the byte-identical derivation to the equivalent attribute path, so a run kept in a file still shares the same store paths. All of this was to simply evaluate the Nix expression. Now we have to build it. Although Nix is great at building derivations in parallel, the recursion here is tail-recursive and therefore serial. I benchmarked the build time of a growing list of button presses and the cost is also linear, as we would expect, with the number of presses. The cost per press is roughly 1.27 seconds with substituters enabled and 0.28 seconds with them disabled. The round-trips cost for checking whether the derivation is in the cache costs noticeably more than emulating the frames does. 3 We’re used to the attribute path being a name , simply a coordinate into a catalogue of things that exist. Laziness means it’s really a program : a sequence of steps the evaluator walks, generating whatever it needs as it goes. Nixpkgs happens to use that machinery to describe software, but nothing about it requires that the tree be a catalogue at all. Coupled with the fact that the store turns out to be a decent persistence layer for reproducible state-machines, makes a our “package manager” reasonable to use for playing Mario. 🍄 The prefix is a precanned sequence of button presses that gets you to the start of level 1-1.  ↩ A screenshot of the frame is also produced, which is used when we want to stitch a video sequence together.  ↩ We can set or if we want to avoid this cost.  ↩ The prefix is a precanned sequence of button presses that gets you to the start of level 1-1.  ↩ A screenshot of the frame is also produced, which is used when we want to stitch a video sequence together.  ↩ We can set or if we want to avoid this cost.  ↩

0 views
Lea Verou Yesterday

Dark mode toggles: two states are enough

A good two-state toggle can actually express all three data model states. Until recently, if you looked at most websites with a theme toggle [1] , you’d find three options: Light , Dark , and System . Examples of tri-state dark mode toggles. In (LTR) reading direction: Ant Design, Red Hat Design System, Web Awesome, Excalidraw, Taiga, Astro, Hero UI. Thankfully, these days the trend has shifted towards a simpler two-state toggle, but tri-state ones are still incredibly common. Examples of two-state dark mode toggles. In (LTR) reading direction: Vitepress, Material Design, Adobe Spectrum, Radix, ShadCN. The rationale sounds plausible: “System” is a different intent than “Light” or “Dark”! One is a policy ( whatever my OS says, do that ) The other is a value ( dark, forever, I don’t care what my OS says. ) Surely, users should be able to express that intent! Except, real users don’t generally seek out dark mode toggles to express intent for things to stay as they are, they seek them out when things need to change. Think of the user goal when browsing a website (as opposed to a separate Settings page, where three states are fine ). E.g. on a documentation site, they may be there to look something up. On a landing page, they may be trying to evaluate whether the product is suitable for their needs. On a media site, they may be there to read the news. On a graphics app, they want to draw something. One thing is for certain: tweaking the theme is not their primary goal [2] . To get in the mindset of tweaking the theme, something needs to be off . When things look right, users just move on with their actual goal instead of thinking about the theme. The tri-state control is solving a largely imaginary user goal that is extremely rare among real users, and does not justify the additional complication and UX friction of a three-state toggle. Worse, it forces the user to decide between choices that produce no visible difference, breaking the principle of feedback . Yes, tri-state toggles are common . That doesn’t make them good . This essay explains why, and how to do better. One of the most common UX mistakes is designing UI around the underlying data model instead of user goals. Good interfaces abstract away the underlying model and expose a model that aligns with user goals (unless of course these happen to coincide, which is rare). This is exactly the case with tri-state dark mode toggles; exposing all three states is data model leaking into the UI. Yes, there should absolutely be three states in the underlying implementation! But at any given point, one of them is irrelevant to the end-user. Users cannot meaningfully express intent about problems they don’t currently have. A dark mode toggle is a temporary comfort adjustment . When it comes to user goals, there are only two real states: You’re reading in bed, the page is a flashbang, you hit the toggle. You’re on a laptop outside and the dark theme is unreadable in sunlight, you hit the toggle. It’s situational, it’s immediate, and it’s usually about the environment you’re in rather than a considered long-term stance on color schemes. A third state assumes a usage scenario where a user visits a website that looks perfectly fine, and still looks for a dark mode toggle to ensure it can continue to look fine in the future. Users do all sorts of weird things, so I won’t assert that this never happens, but it is not a natural user interaction, fueled by a real user goal. Even the strongest proponents of tri-state toggles I have spoken with either admit they have never done this, or bring up some extremely rare, weird one-off edge cases. I tried to ask on social media ( Bsky , Twitter/X , Mastodon , GitHub ), but no matter how hard I tried to word it well, the question kept getting so misunderstood that the data is too noisy to be useful. Besides people misunderstanding the question and talking about these times where they want to override the OS theme, there were also developers answering about debugging use cases (rather than their actual user behavior), or people talking about that one time they accidentally clicked that option. There was even someone who concluded that because they want the site to inherit the OS setting almost always, they should vote “frequently”! 😵‍💫 That said, I don’t think it matters all that much beyond academic curiosity, as a good two state control can actually express all three states — users just need to apply the override the first time it becomes relevant. One could argue that sure, the third state is not frequently needed, but surely it doesn’t hurt to have it there for the one user that will need it, right? But a more complex UI has a cost. It increases cognitive load for interacting with the control and forces you towards certain UI design decisions. A two-state toggle can be very compact: Just a single icon that switches to another when clicked. Some websites do go that route with a tri-state toggle that cycles through three states. Docusaurus for example: But generally, the ergonomics of that are poorer than for the two state toggle, so it is no surprise it’s rare (Docusaurus was the only example I could find). Some tri-state controls go for three icons side by side, which triples the screen real estate used. Others, in an attempt to balance clarity and real estate, resort to a dropdown: That improves learnability, at the cost of efficiency, as it turns a single click interaction into a two-step process. The actual perceived friction is actually worse than one extra click. Perceived friction is not a pure function of user actions, but also of the mental effort required to make a decision, and larger UI shifts (e.g. opening a dropdown) are more cognitively expensive than smaller ones (e.g. clicking a toggle) as the user needs to perceive and interpret a larger area. Guidance towards using tri-state controls is well meaning, but often based on paring good tri-state controls against poor two-state ones. E.g. in this article by Bramus : Above that many implementations I have seen don’t take the “System” value into account. By omitting this option, the sites will never be able to respond to the system preference again, as they always have an override applied. Indeed, a bad two state toggle is worse than a tri-state one. It makes the system mode unreachable once tweaked, making the selection irreversible and violating the usability principle of user control and freedom . A good two-state should be able to express all three states. The idea is that the underlying model is still three states, but only two are shown at any given time : When you press it for the first time, it toggles to the opposite of what you’re currently seeing, and stores the literal value ( or ). The next time you press it, it toggles back to the system default , and removes the stored value. That last bit is the one many two-state toggles get wrong. Storing a value that happens to match the system preference silently converts a temporary adjustment into a permanent pin with no way out. Another common mistake is being overzealous about removing the stored value when the system preference changes, even if the user has explicitly set an override. This evaluation must only happen at user interaction. This is important because many users have their OS set to automatically switch between light and dark mode based on time of day, and removing the stored value proactively would make it impossible for them to actually pin a theme. If a stored override later happens to coincide with the system preference — because the OS changed, not because the user did anything — you keep it . This looks like an oversight — they’re the same now, why not tidy up? Because tidying up silently downgrades an explicit choice into a default, based on an event the user didn’t cause and can’t see. Here’s a concrete scenario that you can navigate interactively ( view on separate page ): An argument I heard when discussing this was “but if the user selects light when their OS is light, then the OS switches to dark, won’t they get confused that the website did not preserve their choice?” People hypothesizing that other people, who are not them, will get “confused” is a bit of a pet peeve of mine in usability discussions, but let’s entertain it for a moment. Here’s that exact scenario: Remember, this control is entirely tangential to the actual user goal for visiting the website. Even if their intent were to pin light instead of reverting to System (light) , this is something they would only notice once these diverge, i.e. the OS switches to dark. At that point, fixing it is a single click away. It’s such an easy fix, that there is no point in dwelling on it further. It’s not that this never comes up, but making the tradeoff in favor of a tri-state control isn’t justifiable, IMO. A tri-state control introduces permanent UI complexity to prevent a one-time, easily fixable problem . Additionally, color appearance is not just a pure function of color components, but also affected by surroundings and other factors. Even if a website implements only two modes, light mode may look slightly different in a light OS vs a dark OS, so selecting it as an override makes it an informed decision . The title and icon could make the state clearer (e.g. the tooltip saying “Switch back to light (system default)” instead of “Switch to light” or the icon having a small screen icon instead of just a sun or moon). But those would need user testing to validate that they are an actual improvement. My concern is that once you distinguish System (light) from light , it (ironically) could become the thing that primes users to seek a third state that they previously had not considered. Even if there is an ingenious UI that exposes three states at the same time without adding any cognitive load or friction (I have some ideas about what that might look like), I’m unconvinced this is a problem worth solving, and feels a lot like the UX version of premature optimization . Although I spent the whole article arguing against tri-state toggles, there are actually valid use cases for them. These are the two cases I’m aware of, but feel free to recommend more in the comments! This article is primarily geared towards a permanently visible toggle in the header or footer . A setting that lives alongside other settings in a settings panel is a fundamentally different usage scenario: It is no accident that while 2-state toggles are becoming the norm for persistent controls, tri-state is (rightly) king for settings panels. Bluesky’s Appearance settings panel. The tri-state is fine here. Showing the “Dark mode” option below even when it produces no effect, on the other hand… Google Calendar. Love the icons, it would be nice to actually indicate what System currently resolves to. I’m not one to praise post-X Twitter, but having two two-state toggles instead of one tri-state is a very interesting design choice. The UX is not quite there, but if done well, I think it could be the best of both worlds when you have the screen real estate. This entire essay assumes the common case where a website only has two color schemes: light and dark, and there is no difference between light mode in a dark OS vs light mode in a light OS. Vadim Makeev had an interesting idea : color schemes should take the underlying OS setting into account. Light mode should be less bright in a dark OS and dark mode should be less dark in a light OS, to reduce the contrast between the website and the rest of the system. I have not seen many UIs doing this, and CSS does not make it easier ( is very much designed around duality), but if you are actually doing this, you have earned your three states my friend , display them as prominently as you like, none of this applies to you! Edit: I reached out to Vadim to ask if he had seen any UIs following his guidance. Here’s what he had to say: Unfortunately, I haven’t seen any websites using this idea. I would say we’re pretty limited with tools currently to do so. The moment we want to override prefer-color-scheme, the whole light-dark() convenience is falling apart. Yet another problem that CSS functions will solve (nothing preventing us from creating a 2-4 arg version of this ). The dark mode toggle is a nice case study, but the underlying lesson is bigger: Users do not seek out solutions to problems they don’t currently have. The tri-state toggle is the GUI version of low signal-to-noise APIs that ask you to pass dozens of parameters that could have sensible defaults, forcing you to decide on problems you have not encountered and are not relevant. Do not flood users with options that are irrelevant to their current situation. Options that might become relevant in the future, should be surfaced in that future, not pre-emptively. Not every state of your state machine warrants visible UI. Ultimately, everything boils down to the very same principle: Respect user effort. Thanks to Chris Lilley and Jake Archibald for reviewing an earlier version of this draft Unless otherwise noted, this refers to a permanently visible toggle in the header or (rarely) footer, not a theme setting in a separate settings panel. ↩︎ This is about users. Yes, the developers of the site may have a goal of testing the theme, but we optimize UIs for being used , not getting debugged. ↩︎ The website looks ok. The user moves on with their actual goal and doesn’t look for the toggle at all. The website is too bright or too dark to be comfortable. The user wants to fix it. Your OS is in light mode and the site has stored nothing, so the page follows along. Flip the OS control to run this the other way round. You toggle. The target is dark , which is not what the OS says, so the site stores an override. The page goes dark . Your OS switches to dark . The override now matches it but is still kept . Nothing visibly happens, which is correct. Your OS switches back to light . The page stays dark , because the override is still active. You toggle. The target is light , which is what the OS says, so the override is removed . The page follows the OS again. Your turn. Both controls are live and nothing from here on is scripted. Drive them in any order and watch what does — and does not — end up in . Your OS is in light mode and nothing is stored. You toggle to dark , which is stored as an override. You toggle again, meaning to pin light . It matches the OS, so the override is removed — you actually got the system default. Your OS switches to dark and the page follows . Not what you meant! But the fix is a single click: light no longer matches the OS, so this time it is an override, and thus pinned, so this can only happen at most once . The user is already in the mode of making decisions about their future The expectation is not that every setting must produce immediate feedback There is a lot more screen real estate to explain three states. Unless otherwise noted, this refers to a permanently visible toggle in the header or (rarely) footer, not a theme setting in a separate settings panel. ↩︎ This is about users. Yes, the developers of the site may have a goal of testing the theme, but we optimize UIs for being used , not getting debugged. ↩︎

0 views
Xe Iaso Yesterday

SigV4 authentication is surprisingly complicated

SigV4 looks simple: sign a request, check the signature. Then you implement canonicalization, clock skew, and a cache that isn't allowed to hold your key. Tigris is a drop-in replacement for AWS S3 (or GCS, anything S3API compatible). As such, we need to be fully compatible with both the mechanisms and semantics of S3 including the SigV4 authentication protocol . This is the lingua franca of authentication in the object storage landscape; even Google Cloud Storage has a way to enable SigV4 support so you can use existing applications against its object storage service. At first I thought that SigV4 was fairly simple. Clients sign requests, servers do the same work and make sure the result matches. The main sticking point is that the cryptography involved is symmetric cryptography, the kind where both parties need to have the same secrets. This makes some scaling issues weird, but we'll get into that in the future. Note This is only going to be talking about authentication (ensuring the identity of a remote client), not authorization (ensuring the client has the permission to do something). Authorization will come in the future for reasons that will become obvious when you see that post. We basically needed to implement a compiler. That is not a typo. At a high level when a client signs a request with SigV4 you get an access key ID and secret access key. The access key ID is functionally a username and the secret access key is functionally a password. Admins can identify keypairs by the access key ID (without special training or tools) and services use the owner of the access key or policies delegated to that access key to determine what actions that client may take. SigV4 uses HMAC (hash-based Message Authentication Code) and SHA-256 (SHA-2 with a 256 bit hash width) to do authentication by creating salted hashes based on request metadata. In order to send a SigV4 request, clients take the outgoing request, reduce it to a canonicalized form, and sign it with a symmetric key derived from the secret access key, the current date, region of the service, and service name, kinda like this Go code: As an example, let's see what a signed request to a HTTP debugging endpoint looks like on the wire with and without the signature: And when you add the signature with : Note This is not a live keypair, it was specifically crafted for this post. Breaking it down we have two extra headers in the request: On the wire, HTTP/1.1 requests look kinda like this: However the headers could be sent in any order, and changing the order of request headers doesn't result in different requests. Additionally any query string parameters could be formatted in any way a client (or server) could imagine, including the use of semicolons to separate values . All attempts to canonicalise HTTP requests MUST deal with this ambiguity and define their own rules. SigV4 canonical requests are made up of a few parts: For that example request, the canonical form would look like this: As the request has no body, the empty sha256 checksum is put as the body checksum. Note This exact approach requires clients and services to buffer the entire request body before processing it. There is a subset of SigV4 that supports arbitrary-sized bodies without having to buffer the entire request using , which requires extra logic that is way out of scope for now. If you want to learn more, give your favourite AI agent the following prompt: Additionally, when you are doing presigned URL uploads in object storage, you replace the body hash with the fixed string when canonicalizing because you have no way of knowing what data the client will upload or what the SHA256 checksum will be. To make the signature, you take the sha256 checksum of the canonical request and then HMAC it against that derived signing key: And construct the header based on your access key ID, service region, and service name: AWS has made an extension to SigV4 that uses asymmetric cryptography called SigV4a (the "a" means asymmetric). Instead of using symmetric cryptography on both the client and server in ways that means the server needs to either know the client's secret access key (or a value derived from the secret access key), SigV4a uses key derivation functions to derive a cryptographic keypair. Servers authenticating requests fetch the public key from IAM. Only the client and IAM know what the private key is, and that private key is what signs outgoing requests. I'd love to use SigV4a more because it makes adding additional services to the mix (such as a git service ) a lot safer as you can have those additional services exist in different trust domains than the core product. This is the core of how microservices end up happening. However, it's not super widely used even within AWS. The only SigV4a use I can find in Amazon is S3 Express Zones , however they may end up using it in other services I'm just not aware of. When I did my own experimentation with SigV4a (where I was implementing my own IAM server so that I really understood this all at a low level), I had to copy a lot of internal AWS SDK code into my repo in order to get it working. I'll talk about SigV4a some more another time. One of the weaknesses of using signatures for API authentication like this is the problem of replay attacks. When you make a naïve signature of a value, there's no real way to tell when that signature was created. If you sign a request to create a compute instance at time instance t0, it's still technically valid at any other time instance tN. This is why the canonical form of SigV4 requests includes the current date and time: This means that the request was signed on July 15, 2026 at 20:54:32 UTC. Time changes constantly (at least at the rate of one second per second!) and the client has to have a working clock in order for TLS to work. Servers can trivially read the contents of and reject old requests. This means that you don't need to add or store nonce (number used once) values with each request because that doesn't scale . Note A lot of the security of this authentication protocol is predicated on TLS being used to encrypt the authentication headers over the wire. If TLS is not in use or is compromised by administrative policy, you're probably in a very weird exceptional situation that is very wrong in the first place. An easy example is an enterprise network with endpoint manglement software that does deep inspection of every user action. As a side effect of this, you need to set a temporal skew window for validating requests. This window needs to be generous enough to accommodate slow clients, sloppy timekeeping on the client side, highly latent clients, leap seconds , or other exceptional temporal phenomena. In general time synchronization is a surprisingly hard problem , so it's best to just be tolerant of clients in order to make things more robust in practice. AWS uses a temporal skew window of 15 minutes for validating requests. I'm going to use a window of 5 minutes for my API because 300 seconds is a nice round number and I don't have to deal with the same amount of legacy code that AWS does. So all of this SigV4 business had been working really well for Tigris. Then we worked with a few customers who needed a local cache to fully saturate their hungry GPUs. To be fair, Tigris is plenty fast, but the real thing that kills AI training is latency and something that runs locally will always be faster than the cloud. In order to provide that sweet middle spot between making everything rely on the cloud and having everything local, we made TAG , the Tigris Acceleration Gateway. This effectively gives you most of a Tigris region in your own infrastructure. When you connect to TAG, your code uses its existing access keypairs, buckets, and code. You point your code to TAG, you point TAG to Tigris, and then everything is cached for you. But how does TAG authenticate with your code? TAG doesn't have access to all your existing API keys (and to be honest it shouldn't), but it's still able to authenticate them with SigV4 authentication. TAG and the IAM server both implement a signing key proxying feature that lets a client and TAG both prove their identity to Tigris. Once that proof is sent, then TAG gets the intermediate derived signing key and uses that for locally validating requests, kinda like this: The actual implementation in TAG involves some derived AES logic so that the derived signing keys are very much limited to the client that requested it (namely: the AES key is the SHA256 encoded form of the proxy secret access key). One of the weird parts is that the canonical form of the proxied requests differ from the normal SigV4 canonicalization process, namely looking like this: This is signed using the same SigV4 signature process as before but added differently to the request: And then TAG reads the response from Tigris, caches those derived signing keys, and then uses those in the standard SigV4 process to authenticate clients: no round trip to the cloud required. The happy path is exactly what I thought it was. Reduce a request to a canonical form, run four HMACs, compare the result. That part fits in an afternoon. Everything expensive lives in the questions around it. Which bytes count as the request? Whose clock decides that a signature is still good? Who gets to hold the key that proves any of it? Each question has an obvious answer, and each obvious answer is wrong in some specific way you only find by implementing it. That last question is the one that surprised me. I read symmetric cryptography as a hard limit: if the verifier needs your secret, the verifier has to be Tigris. It isn't. SigV4 derives its signing key through a chain of four HMACs, each one scoped tighter than the last: date, then region, then service. Those intermediate values can travel without the secret behind them. TAG rides that. The key it holds stops working when the UTC date rolls over. It covers one region and one service. You can't walk it backwards into a secret access key. We also didn't write any of this, which is its own kind of relief. SigV4 is old, widely deployed, and hammered on by every S3 client in existence. Any compatibility bugs here are ours. The protocol's bugs are everyone's. The place a protocol bends is usually some intermediate value that somebody already designed to be thrown away. If you want a Tigris region in your own datacentre, the Tigris Acceleration Gateway caches your buckets locally and authenticates your existing keypairs with the same SigV4 dance your SDK already speaks. : The fixed string to signal to the server which authentication mechanism is in use. The rest of the string is information about the request signature so the server can properly canonicalize the request. : The date and time (UTC) of the request so the server knows when the request was signed. Servers will use this request date in order to reject old requests to prevent replay attacks . The HTTP method ( , , , , etc.) The URI path of the request ( , etc.) The sorted canonical query string (you must exactly match the server-side canonicalization logic) The signed headers terminated with two newlines The sorted list of signed headers joined by semicolons The SHA256 checksum of the request body : the HTTP Host of client requests (EG: tag.default.svc.cluster.local) : the Tigris keypair used to authenticate TAG itself (must be in the same organization as the client) : the time of the request in unix timestamp format : the hex output of signing the canonical form of the request against TAG's secret access key

0 views
Unsung Yesterday

Xcode’s clever minimap

Minimaps are an interesting UI element because they often feel very exciting – something about things being small, or responsive, or maps just being cool? – but fail to actually be useful. I find minimaps for coding particularly tricky because, at least in my world, code all looks very much the same from far away. Seeing a thumbnailed/ greeked version of it felt just like a more expensive and distracting version of a regular scrollbar, without any benefits. But Xcode does something interesting. It allows you to use to create a sort of a “header” for the minimap anywhere you want: = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/xcodes-clever-minimap/1.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/xcodes-clever-minimap/1.1600w.avif" type="image/avif"> The minimap then shows those headers not greeked, but as human-readable text: I thought this was clever, allowing you to see the bird’s eye view of the code not just in the most obvious visual sense, but also as a sort of “table of contents,” adding so much more utility. This, of course, is technically no longer a “zoomed out view” but then again, it’s not that there is some sort of rule that it has to be. As a matter of fact, quite the opposite; it all reminded me of the famous London subway map by Harry Beck, which also broke the expectations in a similar way. The original, geographically accurate map of London’s tube looked like this: = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/xcodes-clever-minimap/3.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/xcodes-clever-minimap/3.1600w.avif" type="image/avif"> Beck decided to take liberties with the geography and reimagine the map as a diagram to help the travellers, and ever since he’s done so in 1933, many transit maps followed suit: = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/xcodes-clever-minimap/4.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/xcodes-clever-minimap/4.1600w.avif" type="image/avif"> #coding #text editing

0 views
Unsung Yesterday

“It’s unclear how Sopwith escaped to the general public.”

Sopwith is a 1984 videogame made by David L. Clark for the original, seminal IBM PC model 5150 . It sports the distinctive 4-color CGA palette and an equally distinctive PC speaker soundtrack. It’s also one of the oldest videogames still in active development, and I was surprised how enthralled I was learning about it. = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/its-unclear-how-sopwith-escaped-to-the-general-public/1.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/its-unclear-how-sopwith-escaped-to-the-general-public/1.1600w.avif" type="image/avif"> (First of all, you can play Sopwith in a browser . Choose “single player” and then “novice” first for the game to tell you about its unusual keyboard control scheme.) The current maintainer of the effort is Simon Howard. He wrote about Sopwith’s interesting history ; I love appreciate this kind of approachable and caring preservation of obscure titles. The history is worth a read. From that, I learned a fascinating factoid. The game was intended as a demo for networking hardware, and the original author didn’t realize the game was “in circulation” for many years: Intended as a trade-show demo, it’s unclear how Sopwith escaped to the general public. David L. Clark didn’t even discover until around 2000 that it had “gotten out”. Little did he know, Sopwith had been circulating for years in collections of early games for the IBM PC. Only a couple of years after the first version was released, ads were appearing in magazines like PC Magazine advertising Sopwith for sale as part of collections of games for the IBM PC The modern edition started by Howard is called SDL Sopwith (SDL being a cross-platform graphics library ): SDL Sopwith is directly derived from the source code to the original DOS versions, and still includes changelog comments that date all the way back to 1984. What I particularly liked about the contemporary Sopwith is its guiding document/​philosophy page , also worth checking out in full. Here are some choice principles: There is something in all this that I feel a lot of software could learn from – not just vintage games. I appreciated Howard being thoughtful about growing Sopwith without forgetting its roots, but also with understanding that some things have changed since 1984. You could imagine remixing “The goal is to be a great old game rather than a mediocre modern game” to something like: Better be a great focused app than a mediocre sprawling app. Lastly, how did I learn about Sopwith? Howard shared this charming installation visual with me: = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/its-unclear-how-sopwith-escaped-to-the-general-public/2.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/its-unclear-how-sopwith-escaped-to-the-general-public/2.1600w.avif" type="image/avif"> #change management #games #history #software evolution Sopwith has a long history that deserves to be honored and preserved. By default, the game should always play like the original DOS version. That means the gameplay in particular should be the same, without any significant differences. Someone who has just discovered the project should find it to be a delightfully accurate recreation of the game they may have played when they were younger. […] Some new features can be enabled by default, as long as they are subtle, unintrusive, carefully considered and can be turned off. An example is the medals feature. The game will never try to be “something it’s not”. This means that it will always have four color CGA graphics, PC speaker sound effects and a low resolution display. It will never add (for example) hi-res sprites or 3D models, digital sound effects or MP3 music. The goal is to be “a great old game” rather than “a mediocre modern game”. New features should be fun and recognize the comical aspects of the game. Features should be carefully considered before being incorporated, not just added arbitrarily and thoughtlessly.

0 views
Kev Quirk Yesterday

📝 2026-08-05 22:52: I have a habit of giving all our pets nicknames, there's many of them for...

I have a habit of giving all our pets nicknames, there's many of them for each of our dogs. We've Nelly for 4 days and already I have: Nells Bells Nelly Bean Nelly Furtado Man Eater (after the song) Nelly the Elephant It's anyone's guess which will stick. Our other dogs, Tia and Sid, are t-bone and squid (short for Sqidley) respectively. 🤷🏼‍♂️ Thanks for reading this post via RSS. RSS is ace, and so are you. ❤️ You can reply to this post by email , or leave a comment .

0 views
Evan Schwartz Yesterday

Notes from the AI Coding Transition

Like many other software engineers, my coding workflow has changed dramatically since the start of 2026. And like many others, I've felt some mix of awe, grief, frenetic productivity, atrophying skills, and understanding less while shipping more. In this moment where the field is undergoing this rapid shift, I've found it helpful to read others' takes on their processes, what they're doing to keep their brains engaged, and their genuinely mixed feelings. Before writing up my own thoughts, I went back through the relevant essays and blog posts from the last ~7 months to find the ones that resonated with me the most. Below are the posts that I especially liked and lines that stuck out from them, either because they gave me some idea about how I might want to use AI or just because they had a particularly incisive description of our field's situation. (Quotes are exact and the bold text is my added emphasis.) If you've read others that you thought were particularly on point, please send them my way! I didn’t ask for the role of a programmer to be reduced to that of a glorified TSA agent , reviewing code to make sure the AI didn’t smuggle something dangerous into production. If you would like to grieve, I invite you to grieve with me. We are the last of our kind, and those who follow us won’t understand our sorrow. Our craft, as we have practiced it, will end up like some blacksmith’s tool in an archeological dig, a curio for future generations. Even if AI agents produce code that could be easy to understand, the humans involved may have simply lost the plot and may not understand what the program is supposed to do, how their intentions were implemented, or how to possibly change it. Peter Naur reminded us some decades ago that a program is more than its source code. Rather a program is a theory that lives in the minds of the developer(s) capturing what the program does, how developer intentions are implemented, and how the program can be changed over time. Cognitive debt tends not to announce itself through failing builds or subtle bugs after deployment, but rather shows up through a silent loss of shared theory. As generative and agentic AI accelerate development, protecting that shared theory of what the software does and how it can change may matter more for long-term software health than any single metric of speed or output. the sense of psychological ennui leading into existential dread that many software developers are feeling Simon: All of the chess players and the Go players went through this a decade ago and they have come out stronger. The Shen-Tamkin study identified six distinct AI interaction patterns among developers. Three led to poor learning: full delegation, progressive reliance, and outsourcing debugging to AI. Three preserved learning even with full AI access: asking for explanations, posing conceptual questions, and writing code independently while using AI for clarification. The differentiator wasn’t whether developers used AI, it was whether they stayed cognitively engaged. metrics don’t capture what’s happening underneath. The mental fatigue of reviewing code you didn’t write all day. The boredom of babysitting an agent instead of solving problems . The slow, invisible erosion of the hard skills that made you good at this job in the first place. You stop holding the architecture in your head because the agent handles it. You stop thinking through edge cases because the tests pass. You stop wanting to dig deep because it’s easier to prompt and approve. There’s no spark in you anymore. Here is something that gets lost in all the excitement about AI productivity: most software engineers became engineers because they love writing code. Not managing code. Not reviewing code. Not supervising systems that produce code. Writing it. The act of thinking through a problem, designing a solution, and expressing it precisely in a language that makes a machine do exactly what you intended . That is what drew most of us to this profession. It is a creative act, a form of craftsmanship, and for many engineers, the most satisfying part of their day. this is different because it is not asking engineers to learn a new way of doing what they do. It is asking them to stop doing the thing that made them engineers in the first place and become something else entirely. a mid-level backend engineer is now expected to understand product strategy, review AI-generated frontend code they did not write, think about deployment infrastructure, consider security implications of code they cannot fully trace, and maintain a big-picture architectural awareness that used to be someone else’s job. That is not empowerment. That is scope creep without a corresponding increase in compensation, authority, or time . From my experience building and scaling teams in fintech and high-traffic platforms, I can tell you that role expansion without clear boundaries always leads to the same outcome: people try to do everything, nothing gets done with the depth it requires, and burnout follows. Now the only limit is your cognitive endurance. And most people do not know their cognitive limits until they have already blown past them. Set explicit boundaries around role scope. If you are asking engineers to take on product thinking, planning, and risk assessment in addition to their technical work, name it. Define it. Compensate for it. Do not let it happen silently and then wonder why your team is burned out. talk about what you are experiencing. The isolation of feeling like you are the only one struggling with this transition is one of the most damaging aspects of the current moment. You are not the only one. Computer programming is, fundamentally, about two things: I have a hard time imagining a future where knowing how to solve problems with computers and how to control the complexity of those solutions is less valuable than it is today, so I think it will continue to be a viable career even with the advent of AI tools. I try not to use LLMs to generate full solutions that I am going to need to support. Whenever I have Claude do something for me, I feel nothing about the results. It feels like something happens around me, not through me. the default output has no soul. It's correct. It's competent. It's fine . And "fine" is the enemy of everything I care about as a writer and an engineer. find it hard to believe that supervising a set of agents is going to lead to an optimal flow experience, because we are more passive, it doesn’t stretch our abilities in the same way, and it requires far less concentration. Will we find flow elsewhere? Solving problems and delivering value will always be rewarding, but I wonder if the optimal flow experience offered by programming has, for the most part, disappeared forever, and many of us will simply find less enjoyment at work. You realize you can no longer trust the codebase. Worse, you realize that the gazillions of unit, snapshot, and e2e tests you had your clankers write are equally untrustworthy. The only thing that's still a reliable measure of "does this work" is manually testing the product. Congrats, you fucked yourself (and your company). You let them run free, and they are merchants of complexity. They have seen many bad architectural decisions in their training data and throughout their RL training. You have told them to architect your application. Guess what the result is? An immense amount of complexity, an amalgam of terrible cargo cult "industry best practices", that you didn't rein in before it was too late. All of this compounds into an unrecoverable mess of complexity. The exact same mess you find in human-made enterprise codebases. Those arrive at that state because the pain is distributed over a massive amount of people. The individual suffering doesn't pass the threshold of "I need to fix this". The individual might not even have the means to fix things. And organizations have super high pain tolerance. But human-made enterprise codebases take years to get there. The organization slowly evolves along with the complexity in a demented kind of synergy and learns how to deal with it. With agents and a team of 2 humans, you can get to that complexity within weeks. And I would like to suggest that slowing the fuck down is the way to go. Give yourself time to think about what you're actually building and why. Give yourself an opportunity to say, fuck no, we don't need this. Set yourself limits on how much code you let the clanker generate per day, in line with your ability to actually review the code. When people say “taste,” what they actually mean is experience. Pattern recognition built up over years of doing the work. But calling it “taste” instead of “experience” does something subtle and harmful: it makes a learnable skill sound like a gift . Doing tasks manually naturally builds up the context required for the decisions involved later because you have time to process everything along the way and construct your mental model of the project's structure. This process requires more attention and context switching, along with way more decisions per hour. Making constant architectural, big-picture decisions while overseeing the work of a cracked junior dev is fundamentally harder than executing standard programming tasks yourself. Decision fatigue is, in my opinion, the next invisible friction point for developers. The problem is that as the coding agents get more reliable, I’m not reviewing every line of code that they write anymore, even for my production level stuff. But I’m not reviewing that code. And now I’ve got that feeling of guilt: if I haven’t reviewed the code, is it really responsible for me to use this in production? There’s an element of the normalization of deviance here—every time a model turns out to have written the right code without me monitoring it closely there’s a risk that I’ll trust it at the wrong moment in the future and get burned. When you stop fighting with hard problems directly, the mental models fade. You stop building intuition. You start pattern-matching on outputs instead of reasoning from first principles. And the worst part –> you don’t notice it happening. The code still ships. The PR still merges. Everything looks fine until the incident at 2am where you genuinely cannot reason about what the system is doing because you never really had to learn it. There’s a good analogy here from aviation. Pilots trained heavily on autopilot gradually lose the ability to fly manually and this isn’t theoretical, it’s contributed to real crashes. I think judgment is built from a specific loop: you form a view, you commit to it, you see what happens, and you update. That cycle, repeated enough times, is what builds calibration. The problem with AI is that it short-circuits the first step. You skip forming your own view and go straight to evaluating someone else’s. Do that enough and the muscle atrophies and again, not dramatically, just quietly. You become a better reviewer and a worse thinker. I did the software engineering equivalent of forwarding an email with “thoughts?” and then going to lunch . The job is the part where your fucking brain has to be in the room. You paste the issue into the machine before reading it. You accept the explanation before forming your own. You create a PR before even understanding what the problem you’re fixing is (!). You request a PR review before reading the diff. You merge because the checks passed and the reviewer approved it and the whole thing smells like progress. here’s the new hard rule I’m following after this “incident”: if I still can’t explain the change, I can’t ship it . No exceptions. many software engineers labor under a delusion that their job is to be excellent at their craft. Of course, wanting to be an excellent programmer is not a delusion; it is a completely legitimate value to hold, and a legitimate purpose to pursue. It’s just not what you’re paid to do at work. Your job , unfortunately, is producing shareholder value . This delusion has been punctured by the end of ZIRP , and again more recently by the rise of AI coding. Today, the ownership mindset defines the role. Although unintuitive, limiting the amount of work that runs in parallel is actually producing better outcomes and outputs. I believe the idea of WIP limits must be emphasised more strongly than before. moving from building features in parallel to building a single feature end-to-end faster. But for me, prolonged use becomes insidious. It's easy to become lazy and hand over thinking to the machine in looking for the next hit of cognitive offload when coding becomes even a smidge difficult. Why type your search and read half a short blog post to understand the problem when the same keystrokes give you the (possible) answer right there and then. When you ask a person to do something, you don’t expect them back in five minutes saying it’s done and ready for the next task. With an agent, that’s exactly what happens. Done. Next. Done. Next. There’s no breathing space. There’s always a next thing to think about. The work used to have a rhythm to it. You’d struggle, you’d get stuck, you’d finally figure it out, and there was this moment of joy when it clicked. Hours in the code, and then done. Figuring it out was the whole reward. That’s what AI can quietly take from me. Not the joy itself, but the sense that the thing was mine, which is where the joy was coming from all along. It hands me the finished thing, the finished thing works, and somewhere in there, I stop being the person who made it and become the person who approved it. AI didn’t take the joy out of coding, I gave it away. a quieter admission: the work isn’t teaching me much anymore, and it’s stopped being fun. That’s a description of becoming a manager . What AI did was give every engineer a small team of tireless, fast, occasionally-wrong direct reports. And with the team came the manager’s problem. The discomfort engineers are feeling right now isn’t an AI problem. It’s a delegation problem, and delegation is the oldest unsolved problem in our discipline. The good news: it’s not unsolved because nobody tried. Managers have been failing at it and slowly adapting for decades. What is in your control is small and it is everything: where you point your attention, what standard you hold, what you decide not to do, and whether you’re honest about which is which. The whole reason “there is too much” feels like drowning is that we keep trying to exert control over the size of the ocean. You can’t. You can only decide where to swim. I want to be able to explain what the system does without first having to ask a clanker to explain it to me. Present-day models tend to produce code that is too defensive, too complex, too local in its reasoning. They avoid strong invariants. They add fallbacks instead of making bad states impossible. They duplicate code, invent bad abstractions, and paper over unclear design with more machinery. If each iteration adds another small defense, the system slowly becomes less understandable while appearing more robust. we may no longer understand the whole system in the same way. We treat it, we monitor it, we stabilize it, but we do not necessarily comprehend it. Some domains will punish sloppiness and demand trust and responsibility, but a lot of software lives in a world where raw speed, quick experimentation, and vast coverage matter enormously. Better visualizations of changes or orchestration or agents will not restore our understanding. Either we need to find clever ways to jolt the human back into the loop and make the changes of the loops legible long term, or we need to find better ways to compose these ever more complex systems. In the old workflow, the creative process happened mostly in your mind. In the new process, you supervise the creative process that unfolds inside the AI’s internal machinations. Now, let’s put the historical novelist in the position of the software developer. She gets a call from her publisher saying they’ve found a way for her to bring four books to market each year instead of one book every two years. They’ve recruited a bunch of top-notch high school and college students who can each crank out five pages a day of competent writing for dirt cheap. The publisher wants the historical novels to maintain the original writer’s level of excellence, or to at least be close, so they’re retaining her services as an editor. The novelist’s job is now to edit the work of the students, each of whom has been carefully prompted to write pages that should, with a little work, be stitched together into coherent chapters. Anyone who has ever graded the work of high school and college kids knows that this is generally not rewarding work. If you’ve ever had to grade a hundred papers in a week, you know what a grind that is. The novelist, like the software engineer, is no longer deeply engaged in her work. Editing is not creating. You do not give yourself over to your imagination. You do not immerse your mind and feelings in the process of invention. Instead, you’re rooting out problems, trying to clean up clumsy wording and redundant descriptions instead. The flow state is gone. You are now a cog in a larger process that doesn’t really value your creativity or your need to exercise it. Worse still–and I have felt this personally after months of reviewing AI-generated code–your skills drop off sharply. When a new issue arises–a feature to be implemented, or a tricky bug to fix–the idea of wasting several hours on it feels insulting. Why should I dig through all that code when Claude can locate the bug in five minutes and start drafting a fix? But I think that creative people choosing to hand over their most imaginative, flow-state thinking to an army of bots will be a mistake in the long run. The feature gets delivered, but I do not really feel like I built it. Maybe this is just another evolution of our profession and in a few years it will feel completely normal. Or maybe one day we will realize that somewhere along the way we stopped programming and nobody really noticed. “I’m not sure I can do my daily job without Claude” The cost was never writing the code. The cost was owning it. A fix you cannot judge, in code nobody on your team understands, is not maintenance; it’s another spin of the roulette wheel . And when the bug comes back wearing a different hat, who do you escalate to? Your vibe-coded grid has no changelog, no support contract, and no team whose reputation depends on it. AI makes touching the code cheap; it does not make answering for it cheap. we are yet to see “mind blowing” software being churned out showing that it is still hard to build great software purely with agents. Coding using models can take you from 0 to 1 very fast. But what about 1 to 10, 10 to 100? In sufficiently large codebases, everyone operates with an incorrect theory of the program . Like many software tools, LLMs are a double-edged sword: they make it harder to construct a detailed mental theory of the software, but they allow you to build a partial theory quickly and they can help you leverage that partial theory more effectively. This is a complex tradeoff that I’m still thinking about. our field is evolving in an incredible and painful (but also joyful) direction if you control the ideas of your software, looking at the code itself is suboptimal and often pointless. large software projects have never been limited only by how quickly an individual can produce code. They are limited by how well people can coordinate their understanding of the system they are changing. The shared language of a software project is not English or Python but it is the common understanding of what its concepts mean, where the boundaries are, which invariants matter, who owns what, and why the system has the shape it does. Before agents, some of this shared understanding was maintained by friction....Some of it was the process by which your understanding became mine, and by which both of us discovered whether we still agreed about how the system worked. The most important skill in prompting is expertise in the domain you’re prompting for. A good illustration of this is Terence Tao’s conversation with ChatGPT about the recently-discovered counterexample to the Jacobian Conjecture. This is not the same ChatGPT I talk to! I couldn’t get to where Tao gets , even with unlimited tokens to burn. There’s a lot to learn about good prompting from Tao’s conversation. Here are a few observations: So why does software keep getting worse across the board? The bar for “user experience” has kept rising, but everything has become increasingly fragile. LLMs are useful for producing code that meets easily and objectively verifiable acceptance criteria which you provide explicitly I've found this simple instruction to vastly improve LLMs' output: " Never write READMEs, docstrings, or comments. I will write those myself later. And yes, I really mean this." Problem-solving using computers Learning to control complexity while solving these problems Write before you look. Before opening a tool, before asking the model, write down what you think. Not a design doc necessarily, just your current understanding of the problem, your instinct about the solution, where you think the tricky part is. Even a few sentences. This forces you to articulate your reasoning rather than pattern-match on someone else’s output. It’s also surprisingly useful as a diagnostic: if you can’t write anything, you probably don’t understand the problem well enough to evaluate any answer. Form a view before reading the suggestion. When reviewing AI-generated code or design, read it critically with your own opinion already in hand. What would you have done? Where does this differ? Why might the model have gone this direction and is it right? This sounds small but it’s the difference between passive consumption and active evaluation. One builds judgment, the other just builds familiarity with AI output. Separate ownership from authorship....You can own code you didn’t write. You cannot own code you refuse to understand. Those are different statements, and the gap between them is the whole job. Decide what you must understand deeply - then triage the rest without guilt. The discomfort is the job, not a bug in it. Acting on incomplete information, sitting with the unease of not-fully-knowing, and committing anyway - that is judgment. Managers don’t feel more certain than you; they’ve made peace with feeling uncertain and moving regardless.... Keep something you understand deeply.... Track what you’re learning, not just what you’re shipping.... Tao’s messages are very short and to-the-point. He doesn’t respond point-by-point to the model, just to the gist The model outputs are much more concise than when I try and talk to GPT-5.6 Sol about mathematics. By signalling expertise, Tao shunts the model into “talking-to-mathematicians” mode, not “explaining-to-amateurs” mode Tao pushes back when the model’s responses look wrong, but he doesn’t directly contradict; instead, he says things like “this looks more complex than I was hoping for” Tao makes several leaps and suggestions himself. He almost never takes the model’s advice about where to go next

0 views

One-shotting a Raccoon Heist game using Claude Fable 5

Back in 2024 I tweeted screenshots of a game concept generated by GPT-3 and some concept "art" created using DALL-E. Today, on the fourth anniversary of that tweet, I decided to see if Claude Fable 5 (running in Claude Code for web ) could build the entire game from the content of that tweet. It did a pretty good job of it! You can play the game here . Here's the GitHub repo , and a short video demo: Your browser does not support HTML5 video. This is the August 5th, 2022 tweet : My GPT-3 prompt back then was: This was in the text completion days, where you asked GPT-3 to complete the rest of the text. It completed: In "Raccoon Heist", you and your team of thieving raccoons are tasked with pulling off a series of daring heists. From robbing banks to stealing priceless art, no job is too big or too small for your furry crew. You'll need to use your wits and your skills to avoid the police and make a clean getaway with the loot. With exciting gameplay and a charming cast of characters, "Raccoon Heist" is the perfect game for anyone looking for a light-hearted caper. My prompt to DALL-E was short: Today's experiment: can I dump those screenshots into Fable 5 with a prompt telling it to write a game, then leave it to its own devices and get a working game at the end? A frustrating thing about Claude Code for web is that it can be hard to test what it's working on while it's still working. I've been using GitHub Pages to work around that limitation, and found it to work really well. Here's my process: That's all it takes! Within about 30 seconds of each push the latest content will be visible at . If you do this with a private repo, anyone who can guess the name of the repo will be able to view the published content. I don't worry much about this myself. Here's the prompt I gave Fable 5 (written in the notes app on my phone - this entire project was conducted on mobile). I accompanied it with the two images from the original tweet. I didn't make any technology choices. I assumed (correctly) that it would probably use Three.js based on previous experiments. Giving Claude access to an OpenAI key turns out to work really well for filling in gaps in its capabilities - in this case we needed some way to generate images to use as textures. Fable is very good at prompting image generators! I said "Work independently - do not ask me to make any further design decisions" because I wanted to see if it could produce a full, working game without any further input from me. I also said "Commit and push as often as possible so I can preview your work". When you use Claude Code in the Claude iPhone app you give it a GitHub repository and it works in a branch. Telling it to "push as often as possible" means commits start landing in that branch straight away. I like asking for as a bit of added flavor - here's that finished file , and the entry it made when it added the dog: New escalation: from night 3 the yards get a patrolling guard dog — a low-poly brown hound with a spiked red collar and a wagging tail. It wanders between random spots, and within 12 units it catches your scent and tracks you by smell (line of sight is irrelevant — it's all nose, shown by a 👃 over its head and barking). It gives up if you open a 17-unit gap. Getting caught messages are now source-specific: guard / headlights / hound. Verified wander → track → caught with an automated test. You can access the Claude Code shared session , and I also used my claude-code-transcripts tool to export my own HTML version which you can find here . Fable started with an index page, vendored a copy of Three.js, then wrote its own gen_textures.py script ( copy here ). It generated the textures and spot-checked them to make sure they looked OK. The metal.jpg file it generated for the trash can looks like this, though I don't think it was applied exactly right in the game itself: The raccoon, dumpster hideout, and both crew raccoons are now perfectly visible on mobile. Committing this critical fix. It decided to generate a title screen, which it did using this gen_title.py script. Here's the prompt it used for that: And the resulting image (which Claude thought was "gorgeous" ) - though I note that when it's shown on desktop it gets cropped to just the top third without the raccoon! Then my favorite change: it added the dog : And did a round of testing on it using Playwright, including another screenshot . It made a few more tweaks, then declared itself done : Raccoon Heist is built, tested, and pushed — 7 commits on , all verified with automated Playwright playthroughs on desktop, portrait-phone, and landscape-phone viewports. The game: You're a masked raccoon sneaking out of your crew's dumpster hideout into a moonlit cul-de-sac. Grab coins, gems, rings, donuts, and fish (they visibly stack on your back and slow you down), then waddle home and bank them before dawn. Each night raises the goal and the heat: night 1 has one flashlight-sweeping guard, night 2 has two, night 3 adds a scent-tracking hound that ignores line-of-sight entirely. A police cruiser sweeps the street with headlights, homeowners randomly peek out of windows, a seagull periodically dive-bombs your unclaimed loot, pizza slices trigger a 6-second speed FRENZY, and a rare golden TV is worth 120 points if you can haul its weight home. Win ranks: TRASH PANDA ⭐ → CAT BURGLAR ⭐⭐ → MASTER OF DISGUISE ⭐⭐⭐, with your best run saved to localStorage. How it's built: Vendored Three.js (no CDN), procedural low-poly characters, and a procedural WebAudio soundtrack (sneaky walking-bass jazz plus all sound effects — zero audio files). All seven ground/wall textures and the title-screen hero art were generated with gpt-image-2 and committed as static assets, so the deployed game makes no API calls. Mobile gets a dynamic touch joystick (appears wherever you touch) plus a DASH button; desktop uses WASD/arrows and Space. Testing caught two real bugs worth mentioning: the canvas rendered at 2× size on any real phone (a assignment was wiping Three.js's inline sizing — desktop DPR-1 tests masked it completely), and the win screen's star-rating div inherited the title screen's full-screen CSS and silently swallowed every tap on the "next night" button. Both are fixed and regression-tested, along with pickup, banking, chase/caught, dawn, night progression, and lose/retry flows. has the full build log. The game is live on the branch — just serve the repo root and open . It's an impressive starting point, but it's not a good game. You move your raccoon around a yard collecting items - donuts, fish, gold coins, jewels - while avoiding guards with flashlights and, in later levels, a dog. You have a limited carrying capacity, and once that's full you need to drop stuff off at the dumpster. If you pick up a pizza slice you get a temporary speed boost. There are no team mechanics at all - there are two other static raccoons next to the dumpster but they're purely decoration. It gets slightly more challenging as the levels progress - the dog introduced in level 3 is the most interesting new mechanic - but it's very, very easy to beat. It's also pretty boring - each night has a fixed duration and you can collect all of the items and then have nothing else to do while waiting for the dawn. I was impressed by the implementation. It's fully 3D, there are trash cans, the flashlight illumination cones are fun, and it has a reasonably coherent visual style. It works on mobile. The music ("a procedural WebAudio soundtrack (sneaky walking-bass jazz plus all sound effects — zero audio files)" according to Claude) is simple but feels about right. As a finished game project, it's mediocre. As a starting point from a single prompt I think it's very impressive. I've vibe coded up quite a few games now. They've all been deeply disappointing from a gameplay perspective - it turns out designing games that are fun remains a uniquely human trait, and one which requires significantly more skill and experience than either Claude or I can bring to bear. That said, I thoroughly recommend tinkering with game development projects as a way to explore the capabilities of agents. It's a fun, low-risk way to try out new things. If you stick at it long enough you might even produce something that's worth playing! You are only seeing the long-form articles from my blog. Subscribe to /atom/everything/ to get all of my posts, or take a look at my other subscription options . Create a new repository for the project at https://github.com/new - this can be public or private, the trick works equally well for both. Start a Claude Code for web session, in the Claude iPhone or Desktop apps or in the browser at https://claude.ai/code Tell Claude what to work on, and encourage it to commit an page as quickly as possible. This will create a branch with a name like Navigate to the Settings -> Pages area for the repository ( in my case), select "Deploy from a branch", pick the branch name, and hit Save.

0 views

News: Microsoft Disclosures Suggest OpenAI Sales Account For Around 70% Of FY26 AI Revenue, more than 7% of FY26 Revenue

Executive Summary: As I discussed in yesterday's free newsletter , analyst estimates have OpenAI and Anthropic making up over 70% of all AI revenues across Microsoft, Google and Amazon. While some might have disagreed, Bloomberg is now reporting that OpenAI "accounted for more than half, and likely about 70%, of Microsoft's actual AI sales during its most recent fiscal year." Bloomberg's maths is explained as such: The actual disclosure from Microsoft comes from its most-recent earnings: To be clear, "run rate" means a non-specific month multiplied by 12, which means that it's very possible that Microsoft actually made far less than $37 billion, but based on that maths, OpenAI would make up roughly 64.8% - that being said, I think 70% or more is a perfectly-reasonable estimate, if not far, far more. The other incredible fact from these disclosures is that OpenAI's spend and revenue share accounted for 7% of Microsoft's $331.8 billion in FY26 revenue - or around 7.26% to be specific. At this point, it's impossible to argue that Microsoft has spent $270 billion in capital expenditures to prop up a single client, and that its overall AI plays have failed to create any significant revenue growth or opportunities. We are now four years into the AI bubble, and Microsoft has little to show for it other than one very large and very unsustainable company that requires near-infinite resources to keep paying its cloud compute bills. If you liked this piece, you should subscribe to my premium newsletter. It’s $70 a year, or $7 a month, and in return you get a weekly newsletter that’s usually anywhere from 5,000 to 18,000 words, including vast, detailed analyses of  NVIDIA ,  Anthropic and OpenAI’s finances , and  the AI bubble writ large . My Hater's Guides To the  SaaSpocalypse ,  Private Credit  and  Private Equity  are essential to understanding our current financial system, and my guide to how  OpenAI Kills Oracle  pairs nicely with my  Hater's Guide To Oracle, as well as the Hater’s Guide To Oracle (Part 2). Subscribing to premium is both great value and makes it possible to write these large, deeply-researched free pieces every week. On Friday, I’ll publish the second installment of the Hater’s Guide to Nvidia — where I’ll take a look at how the AI bubble transformed the company from a pure hardware player to a purveyor of the financial dark arts.  If you want to get in touch — and especially if you have any juicy information about Anthropic, OpenAI, or any other companies in the AI bubble — hit me up on Signal at ezitron.76. I’m also on IB on The Terminal.  Microsoft disclosures and Bloomberg analyses show that OpenAI's compute spend and revenue share accounted for 70% or more of Microsoft's FY26 AI revenues, and more than 7% of Microsoft's overall FY2026 revenues. Microsoft has spent $261.3 billion in capital expenditures since the beginning of 2022. OpenAI accounted for $24.1 billion of Microsoft's FY2026 revenues.

0 views
Herman's blog Yesterday

Committing to creativity

There's a quote from the book City of Thieves by David Benioff where a character laments the fickleness of talent: Talent must be a fanatical mistress. She's beautiful; when you're with her, people watch you, they notice. But she bangs on your door at odd hours, and she disappears for long stretches, and she has no patience for the rest of your existence; your wife, your children, your friends. She is the most thrilling evening of your week, but some day she will leave you for good. One night, after she's been gone for years, you will see her on the arm of a younger man, and she will pretend not to recognize you. While I love this quote, as well as the book (you should definitely read it), I can't help but disagree with the premise that talent and creativity are divinely bestowed and out of your control. I've read a few autobiographies of prolific authors, and they all state the same thing: the words don't come easily and sometimes need to be dragged, kicking and screaming, into the light. The book Creativity by the winner of the most-difficult-to-pronounce-name-award, Mihaly Csikszentmihalyi, shares this conclusion and suggests that creativity is a creature to be nurtured, given space to grow, and needs commitment and a lack of distractions to thrive. The thesis is clear: relying on motivation is a losing game. Motivation is a feeling, while commitment is a decision that outlives the feeling. I've found this to be true in my own life, both in creative pursuits, and in things like relationships and exercise. It is a rare person who is amped to exercise all the time, and so it requires commitment to stick with it. Similarly, showing up well in your relationships can't depend on fleeting passions, because life is bumpy. The nice thing is that commitments become easier with time. I have near zero resistance to exercising and journalling, and have been doing both almost every day for close to a decade. The first year or two were difficult, but once they were a part of my day, they became absurdly easy to continue doing. Relatedly, when I'm in a good writing routine it is effortless to write. However, if I go a month without writing on my blog (as is currently the case), each word is a struggle. Action begets motivation, not the other way around. Creativity is like a muscle. It requires constant use to grow, and atrophies when neglected. But creativity is transferable, and so, engaging creatively has positive effects in other creative domains. This post was inspired by a comic jam I attended this past weekend. The comics were great, genuinely funny, and everyone had a wholesome time. It also made me realise how long it had been since I'd sketched. Conversely, I speculate that forgoing creative thought will lead to it "leaving you for a younger man". She doesn't leave on a whim, she leaves when you stop showing up. It hurts when I see people "brainstorming with AI". This is essentially offloading creative thought to a machine, forgoing the most human of processes. It honestly makes me sad, and I can't help but feel that collectively we’re losing something beautiful. The machines are coming for so much, but we can’t let them have this. So go do something creative. Anything. Make a zine, paint with some watercolours, come up with a song. And keep doing that, forever.

0 views
Kev Quirk Yesterday

New Pets, Sick Pets, and Hemorrhaging Money

Wooh! The last few weeks have been absolutely mental here in Casa del Quirk. Around 2 months ago we started renovating the stables to make things better for the goats and chickens. We've had an extraction system put in so fresh air is always circulating, as chickens are very dusty and goats don't like it too hot and stuffy. While we were at it, we also had some new fencing and gates put in, as well as better hard standing around the stables. Finally we had the old, rotten windows and doors replaced with lovely new ones that have shutters so we can control the air flow depending on the weather. Our stables with the new windows and doors As you can imagine, this has been a significant expense, but we planned for it. So it's all good, but then... Around a month ago we noticed that our older (and favourite) dog, Tia, had rather smelly breath and her mouth was bleeding when she was playing with her ball. We figured it's just old age and bad teeth, so we took her to the vet to have her mouth checked and properly cleaned. Unfortunately it was bad news. Tia has a benign but aggressive tumour growing in her gum. Apparently they're relatively common in older dogs (she's nearly 14) but if left untreated it would get very painful and we ultimately end up having to have her euthanised as it would significantly effect her quality of life. So we paid the unexpected £1,000 (~$1,350) bill to have the initial tumour removed, her teeth cleaned, and a biopsy. In the meantime we were also referred to a specialist to hopefully get the issue sorted properly. Last week Tia went to the specialist to have a full body CT scan. This would give them more info about the tumour in her mouth and what our options are, but also (because she's so old) let us know if there's anything more malicious going on in her body. We have to be pragmatic about this - the treatment was going to be expensive, and we had to make a call on whether we wanted it doing if they found something like cancer lingering somewhere in her body. Luckily for us the old girl came back with a clean bill of health and the tumour was treatable. That was another £1,500 (~$2,020) down the drain, but well spent. The procedure to treat the tumour involved taking around 40mm (~1.5") of her lower jaw away. It's a significant operation and the recovery, especially for an older dog, will be difficult. On Monday she had the procedure, it went really well, and yesterday she came home a day earlier than we expected. She's a tough old girl. She's much better today, but still struggling as she's been through a lot. Having a pet that's in pain is the worst - you can't really reassure them like you can a human, so you find yourself hopelessly watching them suffer. It's heartbreaking, but it's for the best. We're still to get the bill for this procedure, but the vet has estimated an eye-watering £4,400 (~$6,000). If anyone is playing along, that's £6,900 (~$9,200) in unexpected vet bills over the last few weeks. She also isn't insured as it's prohibitively expensive for a dog of her age. So we need to foot the entire bill. Anyway, on to more positive news... Last Saturday we welcomed the newest addition to our family, Nelly the 8 week old Labrador/cocker spaniel cross. The breed is known as a Cockador apparently - very unfortunate name if you ask me. Nelly on her first visit to the vet, being very good. Like any puppy, Nelly is a lot of work. Especially if you add into the mix a very sick dog, and the kids being off school for the summer holidays. I've also been working long hours recently, so my poor wife has been juggling most of this on her own. She's at her whits end. Nelly, Tia, and our other dog Sidney are all getting on well though. Nelly is even helping distract Tia from the pain she's in. It's fair to say that all our savings are gone, and the last of Tia's treatment will need to be thrown on a credit card so we can pay it off over the next few months. Things are extremely busy and stressful at the moment. But hopefully we will have Tia for a couple more years so her, Nelly, and Sid will become a proper little pack that gives us lots of lovely memories. Pets are a gift and we wouldn't have it any other way. I just wish vets bills were a little cheaper! Thanks for reading this post via RSS. RSS is ace, and so are you. ❤️ You can reply to this post by email , or leave a comment .

0 views
Stratechery Yesterday

Google Earnings, The Frontier Case, Amazon Earnings

Google's earnings seemed to confirm the Anthropic hedge; it was Andy Jassy who explained why their — and Amazon's — capex was justifiable.

0 views

Its not nostaligia

“You’re just nostalgic!” They will say when you point out the fact that everything today is crumbling around us. Where every McDonalds has transformed into a grey box without a playplace - because the playplace would mean people stay longer without paying more money, and the grey box can be sold in the future when divestment is on the table. We live in a world where the food that we eat is less nutrient dense than it was 20 years ago - and a 1/3rd the size of what it used to be, while being 3x as expensive. That the men all have half of the sperm count and testosterone they did even 30 years ago. That the car that you drive is built far worse than it was in the 90s and spying on you - because it wasn’t enough to sell you the car and the maintenance - they have to sell your data, too. That the house you live in is built with garbage materials, the corners cut so that the builder can achieve 0.025% more margin. The home is meant to get to the day after the builder’s warranty expires, not even make it to the end of the decade. That in a world of “infinite streaming”, there is nothing to watch. Where the technology you use is continually selling out your privacy and sanity, while you are not permitted to own any of the media you supposedly “buy”. Where everything is locked behind a log-in screen or a paywall. The phone you buy is manufactured to die in 2.75 years so that you must buy another. But not just your phone - your furnace, your car, your walls, your oven, your washing machine and dryer, your computer, everything is meant to /break/ - because if it lasted, you wouldn’t spend more money. Every idea just being re-hashed for corporate profit, nothing new has been said in 20 years. That the things that you have written or recorded are now training data for “AI”. The books that you could be reading are being bought up by “AI” companies, only to destroy them after ingesting them . “But we live longer!” The bugman proclaims as if the length of life is the only metric by which we can gauge life. “But there’s so many things to do!” He proclaims as he eats the slop, consooms the media, and votes for the politician he is supposed to vote for. Everything is getting more pricey, yet you get less every year. Nothing works well, and yet we are called nostalgic when we long for a time that at least we owned the things we paid money for. You essentially have to build everything yourself nowadays if you want any semblance of quality and longevity. The only way forward is radical self reliance. I refuse to pay for things that don’t last, that are meant to break on me, that are meant to trap me into an ecosystem I cannot escape. No - it’s not nostalgia. Things just really are worse. As always, God bless, and until next time. If you enjoyed this post, consider Supporting my work , Checking out my book , Working with me , or sending me an Email to tell me what you think.

0 views
Simon Willison 2 days ago

New release of LLM adds support for reasoning traces, OpenAI Responses, server-side tools, and smarter logging

I released LLM 0.32 this morning, the most significant new version of LLM since the initial launch of the project. The new version includes support for visible reasoning traces, server-side provider tools, redesigned content-addressable SQLite logs, new models, and new features enabled by the OpenAI Responses API. I also released a new version of the llm-anthropic plugin with substantial updates of its own. Running LLM against reasoning models now displays their reasoning traces to standard error, so you can see what they are "thinking" without that information being included in the standard output that you might pipe to another tool. Add to turn this off. LLM includes support out-of-the-box for the GPT-5.6 model family , and the new default model used with is now the inexpensive but capable GPT-5.6 Luna . LLM calls can now use server-side tools from various providers. OpenAI provide a code execution environment as a server-side tool; LLM can now run prompts that benefit from that like so: OpenAI also gets a WebSearch tool. The llm-anthropic plugin adds WebSearch , WebFetch , CodeExecution , and AnthropicMCP , which looks like this: That causes Anthropic to execute MCP calls against my new datasette-mcp plugin as part of a single request/response interaction with their API. The new llm openai endpoint command provides a tool for executing prompts against any OpenAI compatible endpoint as a one-liner. These aren't logged, which makes this a handy tool for running one-off prompts against anything that speaks the lingua franca of the LLM API world. Here's how I use that to run prompts against Gemma 4 12B running in my localhost LM Studio API, via (no LLM installation required) and mixing in the llm-tools-quickjs tool plugin for good measure: LLM's Python API previously required you to create a conversation and then send messages to it one at a time. This was an abstraction over the true nature of LLMs, where each request carries a complete history of the messages that came before it. That abstraction started to get in the way for some more advanced cases, so the new release introduces a parameter that can be used like this: LLM previously returned an iterable sequence of strings from each prompt. This worked great when models returned a string response, but failed to predict the weird shape that models would evolve towards. Today many models return a mix of reasoning text, output strings, tool calls, and even image attachments. With LLM 0.32 you can do this instead : Combine these features and we can finally provide a robust implementation of the semi-standard OpenAI chat completions API, which I've now released as the llm-chat-completions-server plugin: Now you can run prompts against LLM via that server, using the new command! The bigger challenge with that kind of API concerns logging. If we're going to support the pattern where the message sequence is appended to on every request, ideally we can avoid logging all of that duplicate JSON for every turn. The solution is the new content-addressable message store , modeled after Git. You can see the new schema for that in the documentation , but the and commands have both been upgraded to convert that format back into something that's easy to consume. There is a whole lot more in this release. The 0.32 release notes are pretty comprehensive, and the notes for 0.32rc2 , 0.32rc , 0.32a3 , 0.32a2 , and 0.32a0 should fill in any gaps. Existing LLM plugins should all continue to work, but plugins that provide extra models will need to be upgraded to 0.32 in order to participate fully in the new streaming events system. There's a guide to implementing plugins with Structured messages and streaming events in the documentation. I've updated some of my own plugins: Quite a few of the lower-level tools changes in this release were driven by the needs of Datasette Agent . When I started work on LLM, the term "agent" had such a vague definition that I refused to use it. In September 2025 I came around to the idea that " An LLM agent runs tools in a loop to achieve a goal " is well established enough now that I could stop avoiding the term entirely. Tool chains can now pause for human approval and resume from a stored message history - both needed by Datasette Agent. Looking at LLM today it's beginning to look very agent-shaped to me. There's something neat about having a CLI utility that can mix and match different tools from different sources with different models all as a one-liner, and that includes a Python library powerful enough to build systems like Datasette Agent and llm-coding-agent . Maybe the next version of LLM will bake the concept of an "agent" into the core library. I'm still trying to figure out what that would look like. You are only seeing the long-form articles from my blog. Subscribe to /atom/everything/ to get all of my posts, or take a look at my other subscription options . llm-anthropic 0.26 adds support for the Claude 5 family of models, plus , , , and server-side tools. llm-gemini and llm-openrouter and llm-mistral are nearly there, releases coming soon.

0 views
Gabe Mays 2 days ago

4 year follow-up on buying pandemic stock dip + AI reallocation

This is my 4-year investment update following buying the dip on ‘pandemic stocks’ that declined (70%+) in 2022, then reallocating into AI stocks in 2023. I started sharing public updates once a year. Data in this update is as of June 2026. This will be a relatively short update since my thesis is relatively unchanged. See my past updates for more context: Below is…

0 views

Purely functional digital circuit simulator (SICP 3.3)

I have a copy of SICP, or as it is also known, The Wizard Book . This book is widely praised, but I can’t take the time to work my way through all of it. Instead, I’m going to occasionally jump into the parts of it that look interesting. Since last week, are in the process of simulating a digital circuit. The reason this is interesting is the solution in SICP uses hidden mutable state and message-passing to make the code object-oriented. It even uses a mutable global variable for scheduling! We managed to replicate all of that in Haskell, but now we want to refactor the solution to be easier to work with. If we are going to manage the simulation in a pure functional manner, we still have to contend with the fact that during simulation, wires are objects with a fixed identity. A wire does not become a different wire just because its signal changes – the same wire is still hooked up to the same gates. Wires need to maintain their identity somehow. (Continue reading the full article on the web.)

0 views