Posts in Html (20 found)
マリウス 3 days ago

GL.iNet Mudi 7

tl;dr: After almost seven years my Netgear Nighthawk M2 has started rebooting on its own, reporting nonsensical battery percentages and ignoring most of my presses on its touch buttons, so I spent the past three months replacing it with the GL.iNet Mudi 7 ( GL-E5800 ), a 5G NR Sub-6 travel router with two nano-SIM slots plus an onboard eSIM, Wi-Fi 7, a 2.5 GbE port, two USB-C ports and a removable 5380 mAh battery. It is the most capable mobile router I have owned, its 13.5-hour battery rating is close to what I measure, and the LTE reception alone is a clear upgrade over the M2 . Sadly the Tri-band on the box means two bands at a time, there is no MLO at all, both SIM trays are underneath the battery, the touchscreen still can’t get you through a captive portal, and firmware 4.8.5 has a cellular defect that leaves the device on Connecting… after a carrier deactivates an idle data session. If you came to the Mudi line for blue-merle and IMEI randomization, you might be disappointed to learn that this sadly seems to have ended with the GL-E750 . Earlier this year I reviewed the GL.iNet Slate 7 ( GL-BE3600 ), the Wi-Fi 7 travel router that replaced my long-running Linksys WRT3200 ACM as the router in my travel setup . I mentioned in that post that I was also in the process of replacing my even older Netgear Nighthawk M2 , the LTE-A Cat. 20 hotspot that has handled my mobile data for almost seven years now. The M2 has been a reliable piece of equipment, however, it has started misbehaving so badly that I no longer trust on the road. Random reboots, increasingly nonsensical battery percentages, and touch buttons that no longer register most presses make it a tedious device to use, and with it well past any expectation of longevity, I figured it was time to give its successor a proper, multi-month trial before the M2 gives up entirely in the middle of some airport lounge. The device I settled on is the GL.iNet Mudi 7 ( GL-E5800 ), a 5G NR Sub-6 Tri-band Wi-Fi 7 travel router that GL.iNet unveiled at CES 2026 and started shipping back in April. On paper the device is an upgrade over both the Netgear M2 and the Mudi V2 aka GL-E750V2 , which was still a 4G/LTE Cat. 6 device with a 0.96" OLED. The Mudi 7 packs Qualcomm ’s Dragonwing MBB Gen 3 platform, a Wi-Fi 7 PHY with a 6 GHz radio, two nano-SIM slots plus an onboard eSIM, two USB-C ports, a 2.5 GbE Ethernet port, a 2.8" color touchscreen, and a removable 5380 mAh battery, all in a 157x75x22.8mm, 300g enclosure that runs OpenWrt with GL.iNet ’s firmware layer on top. At $419.99, or roughly €425, it is also the most expensive device GL.iNet sells. Just like the Slate 7 , the Mudi 7 is above most consumer travel routers. It comes with a 5G NR Sub-6 Rel-17 NSA/SA modem with LTE Cat. 20 (DL) / Cat. 18 (UL) fallback, and the exact specifications of the hardware are as follows: Apart from having a modem, the second difference from the Slate 7 is the 6 GHz radio, which the Slate 7 lacks entirely. However, the Tri-band on the box is a bit misleading. The Mudi 7 has radios for all three bands, but the chipset cannot drive 5 GHz and 6 GHz simultaneously, so you configure the device as either 2.4 + 5 GHz or 2.4 + 6 GHz. This also means the Mudi 7 has no Multi-Link Operation at all. On the Slate 7 I complained that GL.iNet ’s MLO documentation advertises a 6 GHz band that the hardware doesn’t have. On the Mudi 7 the 6 GHz band is present and MLO is gone, which is an odd trade for a device that costs nearly three times as much. There are two regional variants, GL-E5800NA for North America and GL-E5800EU for Europe, with different 5G NR and LTE band coverage, which is important to travelers like myself. Both variants cover n5, n7, n26, n38, n41, n77 and n78. Beyond that they diverge, as the EU model adds n1, n3, n8, n20, n28, n40 and n75, while the NA model adds n2, n12, n14, n25, n30, n48, n66 and n71, plus n13, n29 and n70 in SA mode only. LTE splits the same way, with the EU model on FDD B1, B3, B5, B7, B8, B20, B28 and B32 and TDD B38, B40, B41, B42 and B43, and the NA model on FDD B2, B4, B5, B7, B12, B13, B14, B17, B25, B26, B29, B30, B66 and B71 and TDD B38, B41, B42, B43 and B48. For my use case (almost exclusively APAC/LATAM) the EU variant turned out to be the more sensible choice, but anyone moving frequently between North America and the rest of the world should read both band lists carefully before ordering. To be fair, though, the Nighthawk M2 splits even harder. Netgear ships that device as at least five separate SKUs, and the band list for each one is quite short. The box itself contains the Mudi 7 , the battery pack, a relatively big travel pouch, a USB-C cable, and the paper manual. No external antennas and no power adapter, which I appreciate, given the chargers I already lug around. The headline feature is the modem, which uses the Dragonwing platform, Qualcomm ’s rebranded enterprise and mobile-broadband lineup. In practice the 4.67 Gbps peak figure is, as with virtually all hyped peak numbers, marketing material. Real-world throughput depends primarily on the carrier’s network, the SIM plan, the spectrum allocation, the band combination, and the signal conditions at your specific location. In my own testing I have seen sustained downlink figures in the 600–900 Mbps range on a properly-provisioned 5G network, and significantly less (in the 100–250 Mbps range) on a more typical mixed NSA deployment. What’s more important, though, is the LTE fallback. The modem falls back to LTE Cat. 20 (DL) / Cat. 18 (UL) and is significantly more sensitive than the M2 ’s aging Qualcomm baseband. In the same hotel rooms where my M2 used to show a single LTE bar at best, the Mudi 7 can consistently show two or three, often pulling more usable bandwidth on the same SIM and the same carrier. Lastly, the Mudi 7 has two TS-9 external antenna ports for those of us who care to bolt on a pair of paddle or directional antennas in RV/cabin/dead-zone scenarios. I haven’t bothered to test these, as my use case doesn’t involve any of that. However, these days most people might have almost exclusively converted to Starlink anyway, so the external antennas might not be as much of a selling point as they were ten years ago. The Mudi 7 has two Nano-SIM slots and one onboard eSIM. Both Nano-SIMs and the eSIM are managed via the touchscreen and the web UI. However, it’s important to note that the Dual SIM Dual Standby in this context means dual standby with an asterisk. The onboard eSIM and SIM slot 2 are mutually exclusive and cannot be active at the same time. The eSIM is disabled by default, and the moment you enable it, SIM 2 stops functioning. SIM 1 remains operational either way, and the modem can auto-switch (i.e. fail over) between SIM 1 and whichever of SIM 2 / eSIM is currently active, but you do not get to keep three simultaneously hot profiles. For anyone hoping to keep a local SIM, a regional roaming eSIM, and a home-country SIM in standby together, this is a bit of a disappointment. Failover itself has also been more rigid than I expected. The web UI exposes the auto-switch feature, including data-usage thresholds and signal-loss triggers, but the failover decision-making has been slow in practice. A complete loss of signal usually does cause a switchover within a reasonable amount of time, but more nuanced situations (such as one SIM throttling without any indication, or losing data while still showing connected ) often require a manual nudge. GL.iNet ’s documentation describes far more sophisticated multi-WAN coordination than the SIM-side auto-switch logic delivers. Then again, to be fair, Mwan3 on the Linksys has had similar issues and I guess down detection is just a complicated thing to get right. One caveat is that both Nano-SIM trays are underneath the battery , so putting a card in or taking one out means having the device powered down, prying off the back cover, and pulling the battery out. On a product aimed at people who buy a local SIM on arrival, that is a weird design. Then again, in many cases the device is probably already powered off because you arrived by airplane anyway. Switching between profiles that are already provisioned (either physical-to-physical or physical-to-eSIM) is one of the things the touchscreen handles well, and it doesn’t normally require any detours into the admin UI. Speaking of which, just like the Slate 7 , the Mudi 7 comes with a built-in touch display, though here it is a 2.8" color LCD rather than the much smaller panel on the Slate 7 . The screen shows the usual variety of things, like signal strength and current network type, connected client count, real-time data usage, battery percentage, Wi-Fi details with a QR code for quick joining, and the ability to toggle the VPN, the Wi-Fi, and a couple of other features without opening the admin UI. Firmware upgrades also display a progress bar on the screen, which (as I had complained about with the Linksys ) is a small but welcome quality-of-life feature. The notable thing missing from the touchscreen is captive portal handling. The moment the upstream WAN is a hotel or airport Wi-Fi network with a captive portal in the middle, the touchscreen is useless and you have to reach for a phone, tablet, or laptop, attach to the Mudi 7 , open a browser, and go through the portal manually before the router (and everything behind it) can reach the internet. But to be fair, a 2.8" panel is probably a poor place to render an HTML login form and a keyboard to begin with. The lockscreen with a 4-digit PIN that was introduced on the Slate 7 is also present on the Mudi 7 , which I once again appreciate, given the kind of sensitive information (carrier and SIM details, VPN state, hostnames) that this screen displays. One annoying quirk is the battery percentage reporting. Both the LCD and the web UI will, after a full charge, stay at 100% for the first 1–3 hours of unplugged operation before catching up to reality and dropping rapidly to whatever the actual state of charge is. The underlying kernel fuel-gauge driver does report accurate values (you can confirm this via SSH and ), but from what I can see the MCU layer that drives the LCD and the admin UI applies some smoothing to avoid the device displaying 98–99% immediately after charging. I would much rather see the truth on the screen than a smoothed consumer-friendly approximation, especially on a device whose entire purpose is to be unplugged for long stretches. The Mudi 7 shipped with OpenWrt 23.05.4 ( , Kernel ), with GL.iNet ’s firmware layer on top. The device runs Qualcomm ’s proprietary SDK and binary blobs. The same software-openness caveats that apply to the Slate 7 apply here as well. You get full root SSH access, the configuration tree, and the ability to side-load the LuCI UI if you want, but you’re stuck with GL.iNet ’s firmware for anything that touches the cellular or Wi-Fi 7 silicon. The original Mudi ( GL-E750 ) is the device that blue-merle was written for, the SRLabs package that changes the IMEI via AT commands on the device’s modem, wipes the stored client MAC addresses, and randomizes the BSSID and the WAN MAC address across reboots, and it is a large part of why the Mudi line got its reputation as the privacy-focused travel router in the first place. However, blue-merle supports the GL-E750 and nothing else, and with the 5G modem, the firmware base, and the entire platform having changed underneath it, there is no indication that this is going to change. If IMEI randomization is the reason you were looking at a Mudi specifically, the Mudi 7 does not give you that, at least today. To be fair, the firmware layer is also what makes the device usable out of the box. The Multi-WAN , WireGuard , OpenVPN , Tailscale , AdGuard Home , DNScrypt-proxy2 , Tor , and the modem management features are all preinstalled and reachable via a friendly web UI, which (as I had mentioned in the Slate 7 review) is a substantial step up over the bare vanilla OpenWrt experience on an older router like my WRT3200 ACM . The Mudi 7 supports WireGuard with up to 600 Mbps. I have been running my own WireGuard tunnel on the device, routing the entire LAN through it, and it has kept up with whatever the upstream 5G or LTE connection could deliver. As with the Slate 7 , Tailscale is available, with the same caveats. Basic connectivity works, but anything beyond the default configuration (exit nodes with advanced flags, subnet routing, tagged ACLs, etc.) is going to require manual intervention via SSH. The Mudi 7 can, like the Slate 7 , run a Tor node and route LAN traffic over it. The moment Tor is enabled, VPNs , DNS , AdGuard Home and IPv6 will not work properly anymore, because the firmware doesn’t (yet) compose these services the way a hand-rolled OpenWrt setup can. Note: As I had explained in the Slate 7 review , these limitations are 100% a GL.iNet issue and not caused by OpenWrt . The same combinations work fine if you wire them up by hand on top of a vanilla OpenWrt installation, including DNS lookups via Tor through DNScrypt-proxy2 . The UI just isn’t there yet on the GL.iNet side. AdGuard Home is, as on the Slate 7 , part of the default installation and just as plug-’n-play. I still don’t use it personally, but the web UI is identical to the one on the Slate 7 and works fine in the configurations I have tested. The Mudi 7 differentiates itself from most travel routers in the number of uplinks it can hold at once, as the device supports up to five concurrent WAN inputs: The cellular modem, the 2.5 GbE Ethernet port (when configured as WAN), Wi-Fi-as-WAN (i.e. repeater mode), USB-C tethering from a phone or a secondary modem, and USB-C-attached USB Ethernet adapters. The firmware uses Multi-WAN underneath, with a friendly UI on top. Router, access point and extender modes are all supported, WDS is not. The device features dual USB-C, with one of the USB-C ports being power-only. The other USB-C port is a fully-featured 10 Gbps port with USB tethering, and USB OTG support. It’s possible to charge the Mudi 7 on one port while simultaneously tethering on the other. USB tethering itself, just like on the Slate 7 , is a matter of a few clicks in the UI. Plug a phone in, enable tethering on the phone, and the Mudi 7 picks it up as a USB Ethernet WAN. The same applies to a USB-to-Ethernet adapter, should you ever need to add a second wired WAN or to bridge into a hotel’s wired LAN where Wi-Fi is unreliable. I have had the Mudi 7 for roughly three months now, and the tl;dr is that the device is pretty solid overall, with a handful of caveats around firmware quirks and the chunkier footprint. Battery life is a bit of a mixed bag here, because it depends a lot on what features/services are running on the Mudi 7 , on the amount of WiFi clients and how cellular coverage is. Let me therefore put it this way: For the amount of features you get with the Mudi , especially compared to my older M2 , the battery life is decent. Having that said, however, I do believe that the Nighthawk , at least in its earlier days, was able to survive longer on a single charge than the Mudi is able to right now. Obviously I don’t have scientific benchmarks to prove it, but I remember vividly being out and about with the M2 for a full day and going to bed with the device only around halfway drained. This is something that I don’t think is possible with the GL.iNet . While the device easily gets through a regular workday, I probably wouldn’t trust it to survive a full day road trip with four friends through a mountainous region. Ultimately, its battery life can be extended using an external powerbank, but that’s clearly not ideal with a device that already weighs 300g on its own. If we’re being honest here, 300g equals about two Google Pixel 5 or two Motorola Edge 30 phones, which can both provide you with a 5G hotspot and which will probably (combined) outlast the Mudi by at least a few hours. So if the pure 5G hotspotting capability is all you care about, the GL.iNet is definitely not a good option with regard to battery life. If, however, you’re looking at it as the centerpiece of your mobile LAN, that will allow you to leave your Slate 7 at home because it supports pretty much every important feature and offers integrated 5G connectivity on top of that, then its battery life isn’t too bad after all. The chassis warms up noticeably under sustained 5G load (especially with a VPN), but never to the point where I’d be concerned about throttling or comfort. The back gets warm to the touch, but no warmer than a mid-range phone under similar load, and certainly not as warm as my old M2 would get at times. Unlike with the Netgear , I haven’t experienced any heat warnings with the Mudi so far. The build quality is solid. The chassis has a reassuring density to it, the touchscreen is responsive, and the front button doesn’t feel flimsy. The back panel is a bit of a weak point, because it is a plastic snap-fit cover protecting the battery and it creaks under pressure. Given that this cover has to be pried off to swap the battery or a SIM, I’m half-expecting it to wear out relatively quickly. Weight and footprint, as I had anticipated in the travel desk write-up , are clearly worse than the M2 ’s. The Mudi 7 is heavier (300g vs the M2 ’s 240g) and noticeably chunkier in both length and width. In absolute terms this is still a small device, but on a packed desk and in a packed bag, the difference is noticeable. The included travel pouch is also larger than the router needs, because most of the extra volume is set aside for accessories. Most people probably won’t use the travel pouch for travel, but rather for storage at home. Charging behavior has been predictable. The 24W PD fast-charging input gets the 5380 mAh battery from 0% to ~80% in roughly an hour, and to full in about an hour and 45 minutes. The device accepts whatever USB-C PD source there is around, including my UGREEN 100W and the Sharge Pouch Mini P2 power bank. If you’re considering this device as a permanent member (or even a centerpiece) of your LAN, I have some good news for you: The Mudi 7 can be operated via USB-C, without its battery plugged in. I don’t know whether this is officially supported by GL.iNet , because when you connect a charger the display will show a battery icon with an exclamation mark inside of it, but long-pressing the front button will turn the device on nevertheless. I haven’t experienced any peaks in power-consumption that would lead to arbitrary restarts without the battery plugged-in, but your mileage may vary. Reliability has been pretty good, and I haven’t experienced any crashes, random reboots, or other issues. The only firmware-level oddities I have encountered are the battery reporting discussed above and the cellular issue in firmware 4.8.5 mentioned in the tl;dr : After a carrier deactivates an idle data session, the router can remain on Connecting… until I intervene. The Mudi 7 is probably one of the most capable travel-friendly mobile routers I have ever owned, and it has a permanent place in my travel setup . The 5G modem, the dual-SIM-plus-eSIM configuration, the dual USB-C ports, the 2.5 GbE port, the removable 5380 mAh battery, and the Wi-Fi 7 PHY can replace the M2 + Slate 7 combo for me, while also covering scenarios (5G, multi-SIM, multi-WAN, USB-C-tethered secondary modems) that the combo never could. While the 300g weight and the bulkier footprint are a step back compared to the M2 , if I account for the added size and weight of the Slate 7 that I had to lug around alongside the M2 to make the LAN work for me, then it doesn’t look as bad anymore. Then again, with the M2 + Slate 7 combo I had the flexibility to only bring what’s really needed, which, for e.g. a day trip, would end up being only the M2 . Apart from that, there is the Tri-band situation, with the chipset only driving two bands at once and offering no MLO at all, the SIM-failover logic, which doesn’t work as smoothly as one would expect, the SIM 2 vs eSIM mutual-exclusion, that is mildly annoying, and the battery percentage smoothing, that makes me distrust everything else the device reports. However, none of these are deal-breakers but more like minor inconveniences. The proprietary Qualcomm blob situation is the more concerning part for me, and as with the Slate 7 , the Mudi 7 is OpenWrt only in spirit , and its long-term wireless and cellular stack support depends entirely on GL.iNet and Qualcomm continuing to cooperate. If your priority is true software freedom, this device, like virtually every other 5G-capable router on the market today, is not for you. Neither is it for you if you bought into the Mudi name for blue-merle and IMEI randomization. If your priority is a modern, travel-ready, multi-WAN, multi-SIM 5G router with a solid (if proprietary) firmware on top, the Mudi 7 is, at the time of writing, the best option I’m aware of. I will be sticking with the Mudi 7 as my mobile-data device for the foreseeable future, and the Nighthawk M2 has been retired to the bottom of a drawer after seven years of service. PS: Make sure to check future updates if you’re interested in the long-term experience with the Mudi 7 .

0 views
Simon Willison 4 days ago

One-shotting a Raccoon Heist game using Claude Fable 5

Back in 2024 I tweeted screenshots of a game concept generated by GPT-3 and some concept "art" created using DALL-E. Today, on the fourth anniversary of that tweet, I decided to see if Claude Fable 5 (running in Claude Code for web ) could build the entire game from the content of that tweet. It did a pretty good job of it! You can play the game here . Here's the GitHub repo , and a short video demo: Your browser does not support HTML5 video. This is the August 5th, 2022 tweet : My GPT-3 prompt back then was: This was in the text completion days, where you asked GPT-3 to complete the rest of the text. It completed: In "Raccoon Heist", you and your team of thieving raccoons are tasked with pulling off a series of daring heists. From robbing banks to stealing priceless art, no job is too big or too small for your furry crew. You'll need to use your wits and your skills to avoid the police and make a clean getaway with the loot. With exciting gameplay and a charming cast of characters, "Raccoon Heist" is the perfect game for anyone looking for a light-hearted caper. My prompt to DALL-E was short: Today's experiment: can I dump those screenshots into Fable 5 with a prompt telling it to write a game, then leave it to its own devices and get a working game at the end? A frustrating thing about Claude Code for web is that it can be hard to test what it's working on while it's still working. I've been using GitHub Pages to work around that limitation, and found it to work really well. Here's my process: That's all it takes! Within about 30 seconds of each push the latest content will be visible at . If you do this with a private repo, anyone who can guess the name of the repo will be able to view the published content. I don't worry much about this myself. Here's the prompt I gave Fable 5 (written in the notes app on my phone - this entire project was conducted on mobile). I accompanied it with the two images from the original tweet. I didn't make any technology choices. I assumed (correctly) that it would probably use Three.js based on previous experiments. Giving Claude access to an OpenAI key turns out to work really well for filling in gaps in its capabilities - in this case we needed some way to generate images to use as textures. Fable is very good at prompting image generators! I said "Work independently - do not ask me to make any further design decisions" because I wanted to see if it could produce a full, working game without any further input from me. I also said "Commit and push as often as possible so I can preview your work". When you use Claude Code in the Claude iPhone app you give it a GitHub repository and it works in a branch. Telling it to "push as often as possible" means commits start landing in that branch straight away. I like asking for as a bit of added flavor - here's that finished file , and the entry it made when it added the dog: New escalation: from night 3 the yards get a patrolling guard dog — a low-poly brown hound with a spiked red collar and a wagging tail. It wanders between random spots, and within 12 units it catches your scent and tracks you by smell (line of sight is irrelevant — it's all nose, shown by a 👃 over its head and barking). It gives up if you open a 17-unit gap. Getting caught messages are now source-specific: guard / headlights / hound. Verified wander → track → caught with an automated test. You can access the Claude Code shared session , and I also used my claude-code-transcripts tool to export my own HTML version which you can find here . Fable started with an index page, vendored a copy of Three.js, then wrote its own gen_textures.py script ( copy here ). It generated the textures and spot-checked them to make sure they looked OK. The metal.jpg file it generated for the trash can looks like this, though I don't think it was applied exactly right in the game itself: The raccoon, dumpster hideout, and both crew raccoons are now perfectly visible on mobile. Committing this critical fix. It decided to generate a title screen, which it did using this gen_title.py script. Here's the prompt it used for that: And the resulting image (which Claude thought was "gorgeous" ) - though I note that when it's shown on desktop it gets cropped to just the top third without the raccoon! Then my favorite change: it added the dog : And did a round of testing on it using Playwright, including another screenshot . It made a few more tweaks, then declared itself done : Raccoon Heist is built, tested, and pushed — 7 commits on , all verified with automated Playwright playthroughs on desktop, portrait-phone, and landscape-phone viewports. The game: You're a masked raccoon sneaking out of your crew's dumpster hideout into a moonlit cul-de-sac. Grab coins, gems, rings, donuts, and fish (they visibly stack on your back and slow you down), then waddle home and bank them before dawn. Each night raises the goal and the heat: night 1 has one flashlight-sweeping guard, night 2 has two, night 3 adds a scent-tracking hound that ignores line-of-sight entirely. A police cruiser sweeps the street with headlights, homeowners randomly peek out of windows, a seagull periodically dive-bombs your unclaimed loot, pizza slices trigger a 6-second speed FRENZY, and a rare golden TV is worth 120 points if you can haul its weight home. Win ranks: TRASH PANDA ⭐ → CAT BURGLAR ⭐⭐ → MASTER OF DISGUISE ⭐⭐⭐, with your best run saved to localStorage. How it's built: Vendored Three.js (no CDN), procedural low-poly characters, and a procedural WebAudio soundtrack (sneaky walking-bass jazz plus all sound effects — zero audio files). All seven ground/wall textures and the title-screen hero art were generated with gpt-image-2 and committed as static assets, so the deployed game makes no API calls. Mobile gets a dynamic touch joystick (appears wherever you touch) plus a DASH button; desktop uses WASD/arrows and Space. Testing caught two real bugs worth mentioning: the canvas rendered at 2× size on any real phone (a assignment was wiping Three.js's inline sizing — desktop DPR-1 tests masked it completely), and the win screen's star-rating div inherited the title screen's full-screen CSS and silently swallowed every tap on the "next night" button. Both are fixed and regression-tested, along with pickup, banking, chase/caught, dawn, night progression, and lose/retry flows. has the full build log. The game is live on the branch — just serve the repo root and open . It's an impressive starting point, but it's not a good game. You move your raccoon around a yard collecting items - donuts, fish, gold coins, jewels - while avoiding guards with flashlights and, in later levels, a dog. You have a limited carrying capacity, and once that's full you need to drop stuff off at the dumpster. If you pick up a pizza slice you get a temporary speed boost. There are no team mechanics at all - there are two other static raccoons next to the dumpster but they're purely decoration. It gets slightly more challenging as the levels progress - the dog introduced in level 3 is the most interesting new mechanic - but it's very, very easy to beat. It's also pretty boring - each night has a fixed duration and you can collect all of the items and then have nothing else to do while waiting for the dawn. I was impressed by the implementation. It's fully 3D, there are trash cans, the flashlight illumination cones are fun, and it has a reasonably coherent visual style. It works on mobile. The music ("a procedural WebAudio soundtrack (sneaky walking-bass jazz plus all sound effects — zero audio files)" according to Claude) is simple but feels about right. As a finished game project, it's mediocre. As a starting point from a single prompt I think it's very impressive. I've vibe coded up quite a few games now. They've all been deeply disappointing from a gameplay perspective - it turns out designing games that are fun remains a uniquely human trait, and one which requires significantly more skill and experience than either Claude or I can bring to bear. That said, I thoroughly recommend tinkering with game development projects as a way to explore the capabilities of agents. It's a fun, low-risk way to try out new things. If you stick at it long enough you might even produce something that's worth playing! You are only seeing the long-form articles from my blog. Subscribe to /atom/everything/ to get all of my posts, or take a look at my other subscription options . Create a new repository for the project at https://github.com/new - this can be public or private, the trick works equally well for both. Start a Claude Code for web session, in the Claude iPhone or Desktop apps or in the browser at https://claude.ai/code Tell Claude what to work on, and encourage it to commit an page as quickly as possible. This will create a branch with a name like Navigate to the Settings -> Pages area for the repository ( in my case), select "Deploy from a branch", pick the branch name, and hit Save.

0 views
Nelson Figueroa 1 weeks ago

Setting Up a Time Machine Drive from the Command Line

It’s possible to set up Time Machine drives from the command line. This is way more convenient and becomes scriptable (there’s one GUI checkbox at the end if you encrypt, so the drive can unlock itself). Also, based on my own personal experience, the Time Machine GUI can be unresponsive so the CLI is much better. I’ll be using a 1TB external SSD in this guide. Here’s how you set up a drive. Some of these terminal commands require Full Disk Access. Specifically, the commands we’ll be using later on. Grant your preferred terminal app full disk access in System Settings -> Privacy & Security -> Full Disk Access. Plug in your drive. Unlock it if you have to. Run the following to get some information we’ll need. Here is what you’ll see if the drive is currently being used for Time Machine: Here is what you’ll see if the drive is brand new: Two identifiers matter here: Confirm that you have the correct disk with this command. We know this one is the external drive due to the line. On Apple silicon the internal drive shows and . We’ll need to erase the disk next. The steps vary slightly depending on whether the drive is brand new or is an existing Time Machine drive. New drives usually ship as ExFAT with an MBR partition scheme, so there’s no APFS container yet. We can erase and convert the disk with this command: This results in a drive with a GPT scheme, an EFI partition, and an APFS container with one volume in it (read more on containers vs volumes in APFS: Containers and Volumes ). It does not encrypt the volume, enable ownership, or set the Time Machine role, which are all things we need. So we’ll delete this newly created volume and create a proper one later. Run this again to figure out the identifier: The identifier is in this case. Now use that identifier to delete the volume that was created in the step: That’s it for this section. Skip ahead to the “Create the Volume” section. If the drive is already being used for Time Machine we need to remove the destination (the disk entry in the Time Machine GUI). First, figure out the destination UUID: Then remove the old destination. We’re using so it’ll prompt you for your machine’s password. Verify it’s gone: You can double check that this worked by checking in System Settings -> General -> Time Machine. There should be no backup drive listed. If it’s still there, you can manually remove it in the GUI. Now delete the old volume. The container stays, so there’s no need to repartition the whole disk: If you get an error like: That’s Spotlight. Removing the Time Machine destination makes macOS stop treating the drive as a backup target, so Spotlight starts indexing it like any other volume and holds it open. Turn indexing off for that volume and try again: Now that the external drive has been erased, we need to create an APFS volume on the drive. Decide if you want your backups to be unencrypted or encrypted and follow the corresponding steps. Note that this only worked for me with the flag in the commands. Do not leave it out! If you skip this option, macOS deletes the volume you just created and builds its own in its place when you register the drive in a later step. It has to do with APFS volume roles. The role is for Time Machine backup stores. You can read more about these roles here: How do APFS volume roles work? . Run the following command to create a volume without encryption: Run the following command. It’ll prompt you for a password for your drive. Run this to confirm everything went well. Under you’ll see if it’s an encrypted volume. You’ll see if it’s an unencrypted volume. Two things to note here: Time Machine refuses any destination that doesn’t enforce file ownership. It can’t preserve the UID or GID of what it backs up without file ownership. Volumes created from the command line have it turned off by default. Run the following to enable ownership, replacing the path with your own external drive’s path: Run the following to double check that it worked. should be : “Registering” means telling Time Machine to use this volume as a backup destination. It’s what the GUI’s “Add Backup Disk” button does. The button and the command we’re going to run both write to . Run the following to register your drive: No output means it worked. appends to your destination list rather than replacing it. If you run into this error, try waiting a bit and then try again: Then run these commands to double check everything went well: The output containing confirms that a destination exists and its volume is reachable. Confirm it points at your volume and not a replacement: That UUID should match the one from earlier. If it doesn’t, macOS replaced your volume with one of its own, which is what happens when the flag gets left out. We can add paths we want to exclude from backups through the command line too. There are three kinds of exclusions: fixed-path exclusions, sticky exclusions, and volume exclusions. But for our purposes we only care about fixed-path and sticky exclusions. Here’s an example of adding a fixed-path exclusion: Here’s an example of adding a sticky exclusion (same command without the this time): Check any path to make sure it was added to the exclusions. You should see next to the path (If you see that means no file or directory is there, not that the exclusion didn’t register): To list fixed-path exclusions we need to read them out of the preferences plist: Sticky exclusions don’t appear in the preferences and are stored as an extended attribute on the item: Removing them is similar to adding. We use instead. The flag is still necessary for removing fixed-path exclusions but not for sticky exclusions. Here’s an example of how to remove a fixed-path exclusion: And here’s an example of how to remove a sticky exclusion: Try manually starting a backup through the command line: The command above may look like it’s stuck if your backup takes a while. You can run this in a separate terminal tab/window to monitor its progress: If you get an error like: That just means macOS started a backup automatically. Once it’s done you can verify with: The path in the output confirms that a real backup exists. You can also check the result code: means the backup was successful. Anything else means the last backup failed. This only applies to encrypted drives. From what I can tell, there’s no way to store the Time Machine drive’s passphrase in Apple Keychain using the command line. If you prefer your drive to unlock automatically when it’s plugged into your machine, you’ll need to do the following. Eject the drive (change the path name to your drive’s): Plug it back in. When the password dialog appears, type the passphrase and check “Remember this password.” That’ll save the passphrase in your local keychain so that macOS can unlock the drive automatically next time you plug it in. No need to type in the passphrase every time. Confirm it worked by ejecting and replugging once more. If you aren’t prompted to type in your passphrase, that means it worked. You can double check via the command line too: If is present, that means the drive unlocked and mounted. If you go through this process a few times there’s a good chance you’ll have several Keychain entries for old Time Machine drives. Deleting a volume doesn’t remove its Keychain entry, you’ll have to do this manually. Normally, these Keychain entries point at volumes. But since those volumes were deleted, the entries are pointing at volumes that no longer exist. We can run the following to list all relevant Keychain entries: There’s two entries. To find the one that actually points to a volume, run: So UUID points to a volume. Which means UUID is safe to delete. We can delete it like so: We can then double check that we only have the necessary Keychain entries left: However, this is just for the sake of being tidy. I don’t think having these kinds of entries in Keychain affects macOS negatively in a significant way. — the whole physical disk. We’ll need this later on when running the command. — the APFS container. This is what we’ll need for the command. (A brand new drive won’t have this one yet. It’ll get created when we erase the disk in a later step.) The volume identifier won’t always be , APFS reuses freed slots so yours may be something like or . Write down the volume UUID, we’ll need it later on to verify everything works. Fixed-path exclusions are tied to a path regardless of what is there. Use these exclusions for anything that gets deleted and recreated, like build caches. Sticky exclusions are the default. They’re tied to the item itself. It follows the file if you move it and copies inherit it. Deleting and recreating a directory loses its stickiness. https://support.apple.com/guide/mac-help/back-up-your-mac-with-time-machine-mh35860/mac https://keith.github.io/xcode-man-pages/diskutil.8.html https://keith.github.io/xcode-man-pages/tmutil.8.html https://eclecticlight.co/2024/11/21/how-do-apfs-volume-roles-work/ https://eclecticlight.co/2024/04/02/apfs-containers-and-volumes/ https://eclecticlight.co/2021/10/12/juggling-with-hfs-and-apfs-partitions-and-volumes-a-primer/

0 views
Maurycy 1 weeks ago

You have been mislead about lightbulbs

There's a story that goes something like this: In 1925, lightbulb manufactures secretly colluded to standardized lifespans at 1,000 hours. They would test each other's products to ensure compliance. This is true. At the time, many bulbs lasted longer than one thousand hours. This is true. Therefore, this was done so that people would always need to buy more light bulbs. This is wrong , but it's the type of wrong that cites its sources and hides in the part you'd never think to fact check: The assumption that a longer lasting lightbulb is a good product. In truth, increasing the lifespan of a bulb makes it worse in every other way... but people think they want a long lasting lightbulb: the purpose of standardizing was (primarily) to avoid a race to the bottom. A old-school lightbulb is a rather simple device : A thin tungsten wire (~20 μm) sealed inside a glass envelope to protect it from air. When current is applied, the wire gets white hot and starts glowing. The most important parameter of a lightbulb is how hot that wire gets: this controls the peak emission wavelength (color) and brightness of the lamp. Room temperature objects do emit light (this is how thermal cameras work), but it's at the ~10 μm range instead of the 400 nm - 700 nm light that we can see. In order to put the emission peak in the visible spectrum, the filament would need to run at ~5700 °C ... that is, the temperature of the sun . No metal can survive these conditions: Tungsten melts at "only" 3422 °C. Since it has the highest melting point of any metal, tungsten is the obvious choice for filaments. However, the metal is quite brittle and drawing it into a wire isn't easy. The first commercialized lamps used carbon filaments that were made by charring plant fibers. However, the carbon would evaporate at fairly modest temperatures ~2000 °C. Tantalum filaments were briefly produced during the 1900s, because the metal was easier to draw into a wire than tungsten. These were the first lightbulbs that could actually be left on at night, although they were quickly replaced with tungsten manufacturing improved. There ware also some experiments using zirconium dioxide ceramics, which become conductive when heated. These allowed lamps to operate in air (obviating the need for a vacuum pump and glass seals ), but were limited by its melting point of 2,700 °C. Since any filament must run below its melting point , the peak emission is always in the infrared. This means that only the extreme high-energy edge of the spectrum is useful for illumination, so a small increase in temperature will make a lamp orders of magnitude more efficient. Also, since this increases the average energy of the atoms, the lamp is able produce shorter wavelengths: resulting in a whiter and less depressing glow. The snag is that when a metal is close to its melting point, the atoms are barely holding together: A hot tungsten filament slowly falls apart as the metal crystals slide past each other. Additionally, atoms can evaporate from the surface until there's no wire left. The rate of both of these processes increases with temperature, so there's a fundamental trade off between color/efficiency and lifespan. The lightbulb everyone always cites in the story is hanging in a California fire department. It's been running nearly continuously for over 120 years and racked up over a million hours of operation. Impressive right? What almost no one talks about is that it's hardly even glowing! Photo taken by Wikipedia user Rjaerial Despite nominally being a 60 W lamp, it draws only 4 watts... and is a lot dimmer than you'd expect from a 4 W lamp due to its poor efficiency. There isn't any documentation, but in all likelihood, the bulb was made wrong and ended up having a very high filament resistance. That's why it was sold for as a night light, because it wasn't usable for anything else. Early bulbs (like that one) were handmade , and quite expensive. Because of this, there were universally optimized for long lives. This resulted in light isn't anywhere near white, and a an efficiency that was a tiny fraction of a modern incandescent lamp. (which are also terrible by any objective standards) Once the production process was automated, new bulbs cost pennies, so it made sense to optimize them to work well... because less efficient bulbs cost more money to operate: Going off modern day prices, electricity costs around 0.10 [$/kW*h], so a 60 W lamp will consume 6$ of electricity over a 1,000 hour lifespan. Considering that such a lamp only costs around 3$, installing one that lasts longer but uses more power would be silly. Case in point , despite the cartel only lasting for 14 years, modern (non-halogen, incandescent) bulbs still last for between 500 to 2,500 hours, a range that includes the cartel's 1,000 hour standard. Instead of "making bulbs last longer", manufacturers spent huge amounts of time and money developing entirely new technology: fluorescent and LED lamps. Because these don't use a wire on the very edge of melting, they can be made to work well and last a long time. Of course, specialized lamps have different requirements: In photography, a truly white light is desirable, which leads to specialized "photoflood" bulbs that only last for a few hours. In the other direction, many indicator lamps are designed for 100,000 hours because they are difficult to replace. Ok, but what's with the testing ? If long lived light bulbs are worse products, why would they need a cartel to enforce a the thousand hour limit? Well, it's because people think long lasting bulbs are a good product: Lifespan is something everyone can understand, and has a direct effect on when you will have to go back to the store: if you saw two 60 W bulbs in a store, one claiming to last 400 hours and the other 2,000 hours, you'd probably get the longer lasting one without thinking about it. It's not that efficiency is hard to understand, but most people aren't doing homework before buying lightbulbs... and it doesn't help that bulb packaging uses input power as a proxy for brightness, so the idea that two bulbs both labeled as "40 W" would have a different brightness is rather confusing. As a result, competition was forcing lightbulb makers to produce worse products. To be clear , I'm not defending the Phoebus cartel: they absolutely engaged in price fixing and other anti-consumer practices, and it's difficult to imagine that profit wasn't a factor when deciding the 1,000 hour standard... but by nature, tungsten lamps are consumable items. I guess the the moral here is reality rarely fits into nice stories. Even something so obvious like "products designed to break are bad" often isn't — every manufactured object is the result of hundreds of overlapping compromises, most of which are invisible to the end user. Also , to preempt the orange site, I'm not saying that planned obsolescence doesn't exist. There are plenty of actual cases of products being made hard to repair so they can sell you another one. ... but lightbulbs aren't a good example. https://www.mouser.com/datasheet/3/299/1/T_1_Wire_Terminal.pdf : Indicator lamp datasheet featuring a 100,000 hour rating. https://www.1000bulbs.com/product/67291/STAG-PH213I.html : A photography lamp that lasts 3 hours. (store page) https://www.youtube.com/watch?v=zb7Bs98KmnY : An excellent youtube video on this topic. https://doi.org/10.1063/1.1657874 : Lab tests of tungsten wire evaporation https://doi.org/10.1016/s0016-0032(25)91062-9 : Brightness and color of light as a function of temperature.

0 views
Justin Duke 1 weeks ago

Cursed knowledge

Nick pointed me towards Marcin who pointed me towards immich's list of cursed knowledge the other day, and it has already become a running joke in the Slack. Here is a baker's dozen of Buttondown's own cursed knowledge: 1 Yes, that's the joke. The Python library assigns the device family to every non-Mac desktop browser The HTML attribute only filters what the file-picker dialog shows you; drag-and-drop and clipboard paste bypass it entirely. Safari and Chrome re-serialize quoted CSS custom-property strings differently when you read them back via : Chrome keeps the single quotes, WebKit rewrites them to double quotes. Django emits a — which fails our CI — for any cache key over 250 bytes or containing a space or control character. Python's has no default timeout and will, given the opportunity, wait forever. SPF directives recursively chain DNS lookups against a hard cap of ten — exceed it and you get a , which can fail authentication for all of your mail. Outlook and Hotmail enforce mandatory TLS but serve a certificate chain rooting at DigiCert Global Root CA (G1) — a root that Ubuntu has since removed from its trust store. Django's tests whether the key exists , not whether its value is JSON . does not lock rows in the order you listed them — Postgres locks them in executor scan order, which is a wonderful way to deadlock two queries that both thought they were being careful. A postgres cannot exceed ~1MB. Stripe will send subscription update events for paused subscriptions. The Python library assigns the device family to every non-Mac desktop browser The HTML attribute only filters what the file-picker dialog shows you; drag-and-drop and clipboard paste bypass it entirely. Safari and Chrome re-serialize quoted CSS custom-property strings differently when you read them back via : Chrome keeps the single quotes, WebKit rewrites them to double quotes. Django emits a — which fails our CI — for any cache key over 250 bytes or containing a space or control character. Python's has no default timeout and will, given the opportunity, wait forever. SPF directives recursively chain DNS lookups against a hard cap of ten — exceed it and you get a , which can fail authentication for all of your mail. Outlook and Hotmail enforce mandatory TLS but serve a certificate chain rooting at DigiCert Global Root CA (G1) — a root that Ubuntu has since removed from its trust store. Django's tests whether the key exists , not whether its value is JSON . does not lock rows in the order you listed them — Postgres locks them in executor scan order, which is a wonderful way to deadlock two queries that both thought they were being careful. A postgres cannot exceed ~1MB. Stripe will send subscription update events for paused subscriptions.

0 views

The Difference Between a Button and a Link

Of the three proposals in the Triptych Project , my multi-year odyssey to add a few small-but-powerful features to HTML, the one that generates the most questions is Button Actions . The proposal itself is very straightforward: we want to add the and attributes to the button. Button Actions are such a simple primitive that people often ask why they’re needed. The answer rests on a distinction that web users intuitively understand but rarely have to think about directly: the difference between a button and a link. I added a detailed “Buttons vs Links” section to the proposal, but I think it deserves a blog-style explanation as well, because most of the existing ones miss the mark. Links represent a destination while buttons represent an action . Functionally, this means that links let users control what context they open in, while buttons don’t. Web browsers offer countless affordances for re-contextualizing a link. Clicking or tapping the link will navigate the current page to that destination. Mouse users can middle-click the link to open it in a new tab or hover over the link to see where it goes. Context menus (right-click on desktop, long tap on mobile) have lots of link-specific options. Web users are very familiar with the features that come with links. They know how to open them, copy them, bookmark them, share them with friends, and maintain them in an inadvisable number of browser tabs. The semantics of a link—the notion that they represent an independently-navigable destination—make it possible for browsers to build all these features. The hyperlink predates the invention of the browser tab, but when browsers added tabs, websites didn’t have to do anything to support them; links represented destinations that could be re-contextualized, so browsers could simply invent a new context for them to open in. Every website instantly got upgraded with a huge new feature. Buttons have none of these features. By default, they cannot be middle-clicked, control-clicked, or hovered over for more information. Buttons don’t allow you to copy their the way you can copy the of a link. Their context menus contain no affordances for saving the action or doing it somewhere else. These are not omissions, but deliberate choices based on the button’s semantics: buttons trigger actions inside a specific browsing context ( almost always the current one ). Copying, sharing, bookmarking—these are all features for re-contextualizing the action of a link. Buttons serve a complimentary purpose because they don’t allow for any of that. A common misconception is that links are for navigating the page, while buttons are for everything else. This is incorrect on both counts. Buttons regularly perform navigations. Clicking a logout button navigates the current page to a logged-out one; clicking a “search” button navigates the current page to the query results. Both of these are navigations in the HTML standard . They change the URL, they get logged in the session history, and they load a new page. And links are often used in situations where they don’t trigger navigations. Relative links can jump around the current page; mailto links can open email clients; download links can save a file to your computer. None of these are navigations, but they are all “destinations” that can be opened, saved, and shared in customizable ways. Navigations should be represented as buttons when their action happens in a fixed context that is not available to be re-contextualized (e.g. bookmarked, shared, middle-clicked, etc.). A frequent place this comes up is with forms that let you edit something you’ve already saved, like a comment on a website. When you click “Edit”, the website shows you an editable text area with options like this: Users will easily intuit what each button does: Should “Cancel” be a link? No! Its job is to close the edit view. Not only does making this a link incorrectly communicate its purpose—visually and otherwise—but it saddles the form “control” with lots of features, like bookmarking and middle-clicking, that have incorrect behavior. There are many plausible ways these buttons could be implemented, but none of those implementations should present themselves to the user as a link. With Button Actions, this entire UX could be implemented with just HTML. The first two buttons use existing HTML features, the second two buttons are made possible by Button Actions. (I’m also taking advantage of Triptych’s DELETE support , but you could do the URL method hack without it.) Philosophically, Button Actions create a generic control that can redraw the current context with a network request. Buttons already have the ability to do this with certain limitations; this proposal removes those limitations. Practically, this allows web authors to implement state transitions by navigating to views. Those views might even already exist as standalone destinations, in which case authors can trivially re-use existing routes while representing the action correctly in the UI. This is a great pattern that HTML should encourage! Unfortunately, without Button Actions, erroneously making this button a link is the only way that we have to implement this interface without scripting. This is obviously an anti-pattern, but it’s an anti-pattern supported by major design systems, because buttons lack the ability to do basic navigation without forms. When building a website that works without JavaScript ( a requirement for UK government sites ), links are the only choice. The US Web Design System (USWDS) even contains an official affordance for it: Add to a link and it will look like a button. Making a link look like a button, however, does not make the link behave like a button. USWDS uses JavaScript to implement spacebar activation , but JavaScript can’t do anything about the litany of other behaviors that differentiate buttons from links, like context menus. Links (even those with ) will still look like links in reader mode or other custom views. That’s the fundamental consequence of violating HTML semantics—the page will be broken for some users because authors cannot possibly account for all the different ways that people interact with a web page. The web simply wouldn’t work if they had to. Navigations are the broadest tool that web authors have to control the user experience—HTML just needs to complete the ’s ability to trigger them. Doing so makes the web simpler, safer, and more accessible for all. If you’d like to support the effort, the best way is to like the Button Actions issue on GitHub and share examples of why the proposal would be valuable to you. “Save” updates the comment with whatever is in the “Save Draft” saves the content of the without publishing it “Cancel” closes the editable form “Delete” removes the comment entirely Big shoutout to The Django Software Foundation for their support of this proposal ! I am currently working on an analysis to demonstrate that Button Actions do not introduce any new XSS vulnerabilities to existing web sites. Supporting this proposal doesn’t resolve that issue, but it does demonstrate to WHATWG that web authors have this need and that it’s worth studying. This blog focuses on buttons that trigger GET requests without forms, because that’s where the overlap with links is, but buttons that trigger unsafe requests without forms are also very useful. requests are probably the most common use-case, because they usually don’t require any additional data. One interesting case for buttons that trigger or requests without a form is “likes” on social sites . HackerNews , for instance, uses links for upvotes, which is in wild violation of HTTP semantics. I understand why they do it though: it’s simpler and works without JavaScript. That’s why it’s necessary to make Button Actions not just possible, but convenient. The proposal addresses all the existing workarounds for the lack of this functionality and explains why they’re not sufficient. The big picture goal with Triptych to is to give web authors a simple and semantic way to model a full CRUD lifecycle in HTML, because that’s all the vast majority of web services need to do. All the Triptych Proposals complement each other—Button Actions are even more useful with additional methods and partial page replacement —but I try to make the case for each one in isolation, both as an anti-logrolling mechanism and because they are genuinely useful on their own.

0 views
Ahmad Alfy 2 weeks ago

Testing Google’s “modern-web-guidance” skill against a real React app

LLM-assisted frontend work has a particular failure mode. The model confidently writes code that was best-practice in 2021. It reaches for , hand-rolls a dark-mode toggle with a class on , or disables the submit button to “prevent” invalid input. None of it is wrong exactly. It’s just a few years stale, because the training data is a few years stale and the web platform moves faster than that. Google Chrome’s skill is a direct attempt to fix that. It’s not a linter and it’s not a codegen tool. It’s a search index over a curated set of best-practice guides , meant to be consulted before you write HTML/CSS/client-side JS, so the pattern you reach for is the current one. I wanted to know whether it actually earns its place in the loop. So I pointed it at a real codebase, the React frontend of a project-assessment internal tool I’ve been building, and treated it as an auditor. This is what came back. There’s no magic. It’s two commands over : returns a ranked JSON list. Each hit has an , a , the web , a , and a semantic score. returns the guide as markdown. That’s the whole interface. The intelligence is in (a) the quality of the guides themselves and (b) whether the semantic search puts the right guide in front of you. Everything below is a test of both. The app is a Vite + React 18 questionnaire. You answer about 10 questions, it computes a recommended tech stack client-side, and you can save, label, and annotate assessments. It runs to 42 source files. What matters for this exercise is that it’s form-and-input heavy but has no images and no marketing-page concerns. So the relevant guidance is going to be about forms, inputs, theming, and layout, not LCP hero images. I did a quick inventory first. The tells were immediate: Then I let the skill tell me what to do about each. I searched for . The top hit came back at 0.75 similarity , the highest of the whole session: The app’s current theming is a wall of light-mode hex: The retrieved guide is refreshingly opinionated about what’s mandatory versus optional. The two non-negotiables: That single declaration is the highest-leverage line the audit surfaced. Without it, even a perfectly hand-themed dark palette leaves the native scrollbars, widgets, and the initial paint canvas stuck in light mode. That’s the exact “white flash on load” that makes a dark site feel broken. Beyond the mandatory two lines, the guide shows how to define color tokens with , so each token carries its light and dark value in one place. Applied to this app’s theme file, the change is small. Every hardcoded hex becomes a pair, plus the two mandatory declarations: And updating the is just one line: (The dark values are illustrative inversions. The point is the shape of the change, not the exact palette.) What surprised me is that the guide doesn’t stop at CSS. It carries a section on the design of a theme toggle. This is part of the guide’s own text. You can read it with , or straight on GitHub in the dark-mode guide . Its UX considerations subsection makes the sharpest call, arguing that you should not build the toggle most of us reflexively build: DON’T expose all three states (system, light, dark). … Two of the three options always produce the same visual result, violating the principle of feedback. Instead it argues for a two-state control, “follow the system” and “the opposite of the system.” It also spells out the edge case that trips people up. If a user pins dark and then switches their OS to dark too, the site must stay dark rather than flip. That’s product judgment sitting inside a CSS guide, and it’s exactly the kind of thing a model won’t reliably volunteer on its own. Finally, because is newer than , the guide hands over the fallback so you don’t have to reason it out. You degrade through , then upgrade with where exists. It also ships a copy-paste script to prevent the theme flash for users who have pinned a non-default choice. That script is a plain inline one, deliberately not and not a module, so it reads the saved preference before first paint. And that brings up the skill’s best structural feature. When a guide leans on anything newer than the long-settled web, it keys its browser-support advice to Baseline . For , the dark-mode guide returned that it’s widely available and has been Baseline since 2022-02-03. For , it returned that it’s newly available and Baseline since 2024-05-13. This matters because it turns “should I use this?” from a vibe into a decision rule. The skill’s own instructions say Baseline-Widely-available features are safe to use unfenced, while newer features must carry the fallback the guide provides, unless you’ve declared a custom browser-support policy. In other words it defaults to safe, and it tells you exactly where the risk line is instead of leaving you to guess. is safe to just ship, while gets a -guarded fallback. That’s the correct call, and it made it without me having to ask. I searched for . That surfaced the guide at 0.50 similarity, with and an guide right behind it. Here’s the app’s save surface, lightly trimmed: The guide’s very first rule is blunt about it. “DO use the element to wrap interactive controls… DON’T use for primary submission buttons.” The rename field and the notes editor elsewhere in the app repeat the same -plus- shape. The practical cost of the current approach isn’t abstract. Because there’s no , pressing Enter in the label field does nothing , and that’s a reflex every keyboard user has. The fix is small, and the guide hands it over directly, including the AJAX-friendly submit handler: Wrap the input and button in a , make the button , and Enter-to-submit comes back for free, along with native form semantics for assistive tech. I’ll give the skill credit for a fair grade, too. One thing the app already does right showed up in the same guide. The save button disables itself while a save is in flight, and the guide explicitly blesses that. “DO disable the button after a valid submission is clicked to prevent double-posts.” This is the opposite of the anti-pattern from the intro. Disabling after a valid click to stop double-submits is good, while disabling up front to block an incomplete form is the dead end. A good auditor tells you what to keep, not only what to change. I searched for . It surfaced a cluster of tightly-scoped guides, , , and , all built around and . This is where the guides go deeper than a model’s default answer. Ask a chatbot “how do I validate a form field” and you’ll usually get an handler that yells the moment you type one character. The guide instead ships a timing matrix : The guide even boils it down to a single rule. “Validate on to avoid premature warnings while typing, and reset error states on as soon as the user attempts a correction.” The modern platform gives you this essentially for free via the pseudo-class, which only matches after the user has interacted. The app’s ad-hoc error paragraphs are accessible, which is another thing it got right, but they’re wired by hand where the platform now has a purpose-built primitive. The guide’s section 3 covers , , and , the attributes that tune autofill and the on-screen keyboard. None of the app’s inputs use them, and this is the one finding where the honest answer is a polite no. The questionnaire is almost entirely , where you pick one of a handful of options. Radios don’t take an or an token, because there’s no keyboard to optimise and nothing to autofill. The only free-text fields in the whole app are a “label” and a “notes” box, and neither maps to a standard autofill value. So the guidance is correct in general and largely irrelevant here, and noticing that is the actual work. The tool returns a rule. Deciding it doesn’t apply to a radio-driven form is a judgment call it can’t make for you. This is the clearest example in the whole audit of why the skill is only half the loop. One line from the section does still land universally, though. Text inputs should be or larger, because anything smaller triggers an auto-zoom on iOS Safari the moment the field is focused. The app has in two files. The well-known modern fix is (dynamic viewport height), which accounts for mobile browser chrome that ignores. On iOS Safari, is measured against the expanded viewport, so the bottom of a layout sits behind the address bar. My query returned the broad and guides rather than a laser-focused “use dvh” atom. The right answer is almost certainly inside those guides, but the search didn’t hand me a -titled hit the way it did for . Which is a fair segue into the honest assessment. is not going to catch your bugs and it won’t rewrite your components. What it does is remove the single most common source of stale frontend code, the confident-but-outdated pattern. In one afternoon pointed at a real app, it correctly flagged a missing declaration, a set of forms that skip native submission, a validation approach that predates , and a pile of missing input attributes. For each one it handed over current, Baseline-checked, copy-pasteable guidance, while also telling me which of my existing choices to leave alone. One reframe stuck with me. It’s less a tool you run and more a standard you consult . The best time to reach for it isn’t during a cleanup audit like this one. It’s the moment before you write a component, when the model in the loop (human or AI) is about to reach for the pattern it already knows. Half the time, the pattern it knows is three years old. This is the cheap check that catches it. Zero elements. Every data-entry surface is a bare plus a . No , , or on any input. A hardcoded light theme. defines tokens as literal hex values, with no , no , no dark variant. in two places. Some genuinely good instincts too, like / grouping, on errors, and . The guides are high quality. They read less like scraped blog posts and more like a curated reference assembled by people who live in the web platform, close to the specs and the browser internals, but writing for the developer who actually has to ship. The mandatory/optional split, the timing matrices, and the “don’t build a three-way toggle” UX arguments all read as earned judgment, not a spec dump. Baseline-keyed fallbacks where a feature needs one. This is the single best thing about it. It converts “is this safe?” into a date comparison and provides the exact fallback when the answer is “not yet.” You don’t have to look for the fallback, it comes with the guidance. It grades fairly. In two places it validated code the app already had right. An auditor you can trust to say “keep this” is one you’ll actually keep running. Framework-agnostic by design. Every guide is HTML/CSS/DOM, and adapting the pattern to React was trivial. Nothing assumed a framework, so nothing fought mine. It’s local, self-contained, and keyless. The semantic search runs on your own machine through a small on-device model, so the matching itself makes no network calls and there are no API keys to manage. The npm package ships with no extra dependencies, which keeps latency low and the supply-chain surface small, and the CLI can run fully offline. By default the tool reports anonymous usage statistics to Google, including your search queries and guide retrievals, which you can turn off by setting . It doesn’t read your code. You (or your agent) do. This is the big one, and it’s worth being exact about, because it changes how you run the skill. Nothing in this audit was automatic. The app had to be read, the suspect patterns spotted, each one turned into a search phrase, and the returned guidance compared back against the actual lines. There are two ways to do that. You can drive it by hand, deciding what to search, reading the guides, and applying them yourself. Or you can hand the whole loop to a coding agent, which is what I did here. The agent inventoried the frontend, chose the queries, retrieved the guides, and did the comparison, while the skill only ever answered “here is the current best practice for X .” Either way, the skill supplies the standard and something else supplies the code-reading. Point it at a codebase with no idea what you’re looking for and it hands you nothing back. Semantic search has a recall ceiling. hit at 0.75, but the answer never surfaced as its own result. When a query returns only broad category guides, you have to retrieve a large omnibus guide and read it yourself, which brings up cost. The guides aren’t small, but the skill is upfront about it. Every search result carries a in its JSON. The guide reports about 4,500, and about 7,100. I didn’t measure those myself, because the tool hands them to you before you fetch, so you can weigh the cost. Retrieving a few of them still meaningfully fills a context window. That’s fine for a deliberate audit, but something to watch if you wire it into every edit.

0 views
Simon Willison 2 weeks ago

A Fireside Chat with Cat and Thariq from the Claude Code team

Earlier this month I hosted a fireside chat session at the AI Engineer World's Fair with Cat Wu and Thariq Shihipar from Anthropic's Claude Code team. We talked about Claude Code, Claude Tag, Fable, coding agent security, evals, tool design, and how Anthropic use these tools themselves. The full video of the session is now available on YouTube . Below is an edited copy of the transcript, with extra links and my own bolded highlights. A few top-level notes if you don't want to watch the video or wade through the whole transcript: Simon: Claude Code came out in February of last year — it's under a year and a half old, and it was originally just a bullet point on the Claude Sonnet 3.7 launch . How has what you do on a day-to-day basis changed in the past year , now that we have these coding agents that actually work for us? Cat: I remember when we first came out with Claude Code and Sonnet 3.7, you would give it a task and you would have to closely monitor every single little thing it tried to do. I would read every permission prompt extremely carefully. I would frequently say no — no, no, no, did you check this file? Did you check that file? And now it's been incredible with every model generation. I feel like we've all gotten a chance to take a step back and delegate a lot more of the menial implementation to Claude . It's freed up a lot of our time to think about more creative work, like: what is the right experience that we should be providing to our users, now that we know Claude Code can implement a lot of it? And now with Fable it's a totally different step change improvement. We see for a lot of our use cases that you can actually one-shot a ton of features with Fable now . Thariq: I remember the first text I got about Claude Code. One of my best friends was like, "You need to go try Claude Code." It was about when Opus 4 came out, and I tried it and I was like, "Oh, shit. I need to work at Anthropic now." And that was Opus 4 — great model, but you were reading permission prompts. It's kind of crazy how much amnesia we have, where I'm like, oh, auto mode has always been here, right? I don't even remember pressing yes and allow. For me, the big thing I'm trying to push myself on is that we have to do higher quality work than we've ever done before . The outputs are incredibly high quality. I've been using it to edit videos a bunch , and I'm like, okay, it has to meet the very exacting demands of our brand team in a couple of hours or we just can't do it. That's how I'm trying to shift with Fable: the best work we've ever done, faster than we've ever done it before . Simon: What's a piece of conventional software engineering that was true a year ago that you don't think holds anymore in this new world? Cat: One of the biggest shifts we're seeing in the eng skill set: two years ago it was pretty typical for a product manager to go talk to a bunch of customers, align over the course of six months with cross-functional teams on some PRD, and write a thorough spec on exactly how we'll implement this before the first line of code gets written. Now things are completely turned the opposite way. For a lot of engineers, the push I would give to folks in the room is to develop more of your business sense and product sense on what it is we should build , because the timeline between having an idea and building it is so much shorter — it's down from six to twelve months to maybe even a week. That means all of us need to have better taste on what is worth building, what will actually inflect the businesses we're working on. So it's an increase in value on product taste and business sense , and a bit lower on execution in most product domains. Of course, for infra there's still a very heavy emphasis on making sure all the details are right. Thariq: For me, it's that rewrites are now good . Simon: The worst thing you could do is now actually fine! Thariq: Exactly. All the Mythical Man-Month stuff — never rewrite — I'm pro-rewriting now. If you have a good test suite — and I think the rewrite actually forces you to make sure you have a good test suite — but I think what people undercount is that a codebase is a spec, and maybe it's the only copy of the spec that you have , because no one knows every branching part of the codebase. You can take this as an artifact and distill it or create other versions of it. We rewrote Bun in Rust and it works great — it's live for me right now. Simon: You're not shipping Claude Code on Bun-in-Rust yet, right? Thariq: Internally we have. (Actually it looks like Anthropic started shipping Claude Code on Bun-in-Rust to everyone on June 17th .) Simon: The other big launch recently was Claude Tag — that's what, a week old now, at least for the rest of us. I understand it's being used at Anthropic by non-engineers a great deal. What kind of things are non-engineers doing with Claude Tag? Cat: Claude Tag is a Claude that lives in your team's collaboration tools. We launched it last week within Slack. The thing that's different about Claude Tag is it's multiplayer by default . Once you add Claude Tag to a Slack channel, you can chime in, your teammates can chime in, and you can collaborate together on the PR. The other big difference is that it's proactive instead of reactive. You can tell Claude Tag, "Hey, monitor every bug report in this channel, put up a PR to fix it, and tag the engineer who last touched this part of the codebase," and it'll do it for the lifetime of the channel without you having to manually tag it in. And the third big shift is that we've added team memory into this . If you tell Claude Tag your preferences in the channel, it'll remember them for every future post. If you always want it to debug outages but you don't want it to debug warnings, just tell it that in natural language in the channel and it'll remember it for you and everyone else on your team. Internally, we see Claude Tag as the evolution of Claude Code. We see this as a large shift in how we work internally. Claude Tag currently lands 65% of our product eng PRs. Simon: For all of Anthropic, or just for Claude Code? Cat: This is just for our product engineering team — our internal version of Claude Tag lands 65% of our product PRs right now . And this is a huge shift; this is more than 50% of our PRs. The way we see people split work between Claude Code and Claude Tag is: Claude Code is still the best place for your most complex tasks, when you're interactively iterating with the agent. But Claude Tag is great for having it work proactively on your behalf , so you no longer need to manually kick off Claude Code for all the bug reports that come up for features you're working on. Thariq: And for non-coding cases: for example, before this talk we asked Claude Tag, "Hey, when is Fable releasing?" We wanted to make sure we'd line it up with the announcement. Claude Tag would search our Slack and look at who's been saying what. As a search engine for your company, it's really valuable. It has all the context for your product, so you can ask it metrics-related questions — often when you're making decisions you want them informed by what the metrics say, so you hook it up to your event store. I've seen our marketing team do things like, "Hey, tell me about this feature." They're not programmers, but Claude is a programmer — it can clone the codebase and say, "This is the feature, this is what it looks like, this is a recording of me using the feature ." It enables a whole wide variety of things, and I think we're still early in figuring that out. Simon: One of the problems I've had with coding agents is that I get how to use them as an individual, but I'm not really clear on how to use them in a team environment. It sounds like Claude Tag is your current answer to that team collaborative layer for this stuff. Cat: Exactly. And a large percentage of our sessions are actually multiplayer right now. Maybe I say, "Hey, I think we should implement this new feature in Cowork," and I'll tag in Claude Tag to do a first pass at it. Then I'll tell Claude Tag, "Share a recording of your final implementation," and I'll tag in design to take a look. They'll nudge it, then pass it on to eng to take it to the finish line and get it out to prod. It's been this very fluid experience. We're still trying to iron out what the social dynamics are for steering the same session , but we've found that people just observe how others use it and follow those social norms — it's been pretty intuitive for us to integrate Claude Tag into our teams. Thariq: It's great for teaching people, and also for reducing slop, because the fact that everyone is seeing you use Claude together sort of levels up how you use Claude as well . This reminded me of how Midjourney solved the challenge of teaching people advanced image prompting by enforcing prompting in public in their Discord channels. Something I've found really hard myself is knowing when a feature is worth shipping now that the cost of actually building features has dropped so much. Simon: How do you deal with the hardest problem in all of engineering — prioritization? How do you decide which features are worth building and shipping when building a feature is so much more inexpensive now? Cat: This is the hard thing. There are a few ways we approach it. One is we dogfood our products every single day. Whenever there's something we want to be able to do in our products that we're not able to, instead of finding a different solution we fix our product so it can support that case. We have a very heavy dogfooding culture internally. Before we share our products with everyone in the world, we share them with everyone within Anthropic, and with some early customers who give us very honest feedback about it — the more brutal the better — and we iterate until people love it. We have an internal bar for the number of active users and the amount of retention a feature has to have before we share it with the world. Because this bar is very clear, every engineer knows what they're trying to hit. I think this also levels up our polish, because if the feature isn't polished, people will churn — and then we shouldn't ship that feature. Using internal user-retention to decide if a feature should ship makes a whole lot of sense to me. Simon: Do you have an example of a feature which surprised you? You rolled it out and the engagement was off the charts — something unlikely to be shipped that turned into a real product thing. Cat: I do have one. A lot of folks on our team love remote control . Remote control lets you use your mobile device, or Claude in the web browser, to connect to a local Claude Code session running in your CLI. I never have this need, because I just kick off the task directly on mobile and it runs in a cloud session without using my local environment — I think because I'm doing very easy coding tasks. It was something I didn't totally understand; I was like, hey, people should just set up remote dev environments. But in practice, once we rolled out remote control, so many people I talk to told me that what they do every night is plug their laptop into a power charger, open a bunch of remote control sessions, lock the screen, and then use their mobile phone from their couch to control Claude Code . So this has become a flow we're now leaning into that I didn't originally get — but now I do. One of the over-arching themes of the conference was review: how much attention to people spend to reviewing code written for them by coding agents. I was very keen to hear the Claude Code team's take on this! Simon: How does code review work? Does a human being review every line of production code that makes it into Claude Code? And if not, what are you doing — how do you keep the quality up? Thariq: It varies on the task a lot. For important areas we have code owners. The system prompt is an example where we have a code owner — you really need to get their approval. Simon: So the code owner is directly responsible for the quality of that area of the code. Thariq: That's right. Cat: And they need to approve any PR that touches it. Thariq: We have our code review GitHub bot review everything — that goes on every PR, and often it's doing the bulk of the review. Something I've seen on the team is that for more complex PRs you might make an artifact to explain the PR so that other people can then review. And we invest a lot into verification, CI/CD, things like that, to make sure that any time anything fails we have a test. We have a really robust environment where Claude can control Claude Code and test it. So there's a multi-pronged approach to code review. Cat: In general, we are trying to move to a world where humans don't need to be in the loop . For the most critical changes to the core of Claude Code, and the cores of other products, there is always a code owner and they do manually review all the changes. But increasingly, for the changes at the outer layers, we actually have Claude code review fully review those . That sounds pretty scary, but we've had a six-plus-month-long process to get here, and there are baby steps that you take to build up trust with code review . In the beginning we had human review for everything, and then increasingly we would say, okay, for code changes that touch these files, code review is catching 100% of the issues there — so we actually don't need a human manually reviewing those . And when we have incident review, we look at the PRs that caused the incident and say, okay, how do we update code review to catch that? — and we take those PRs and add them to an eval set to make sure our future changes to code review never regress that metric. Removing humans from the code review loop is a big step forward. It can sound scary, and it's not something you can do overnight, but it is something you can do through many months of investment in the infrastructure to give you the confidence that code review is catching everything you care about. So the key seems to be constantly iterating on the automated review systems themselves, in order to build trust in them over time. We got deep into evals - another hot topic throughout the wider conference. Simon: I know that Opus 4.8, if I ask it to build me a JSON endpoint that runs a SQL query and outputs JSON, is just going to get it right — that's not something I have to review closely. But then a new model comes along and I don't know how to build trust in Fable quickly, that it's not going to mess things up that Opus didn't. How does the new model affect your intuition for what it can do and what it can't do? Cat: The main reason we're building up this eval base over time is so that new models can be a drop-in replacement . When we have a new model, we run the whole eval set and make sure that, for example, Fable is strictly better than Opus 4.8 — and that gives us the confidence to drop it in. Simon: Are those model evals for Anthropic as a whole, or Claude Code team-specific? Cat: We have both. We have evals on our team, and we run code review across every repo within Anthropic, so we have evals for that. And for things like auto mode, we not only have evals across every user within Anthropic — we've also commissioned multiple external testers to red team it, to create environments with prompt injections and malicious inputs, and make sure that auto mode doesn't let any of those pass . Simon: I want to know if the system prompt improvement I made actually improved the product — that's the most basic form of product-specific eval, and I still don't have a great feel for how to do that. Is that something you're doing such that you have complete confidence that a tweak you've made to the system prompt results in better output? Cat: We don't have complete confidence, but we do a lot to make sure that we don't regress performance. The starting point is a suite of external evals that we trust, and we complement that with an even larger suite of internal evals that we trust. To start, we mainly optimize for capability : given a complete definition of a task and the full codebase, does Claude make the right decisions, fully fix the bugs, and pass all the tests? That's the starting point and the thing we optimize for, because it's most directly what users want. But there are a lot of behaviors that impact how users feel when they work with Claude Code. For example, people really don't like it when Claude Code says it's time to go to sleep. Or people really don't like it when it says, "Hey, I finished two out of five parts — do you want me to continue?" Yes, please continue. So we're building up a set of behavioral evals to catch these. And as we get user feedback — please be loud with us about your user feedback — we rank the priority issues and go down one by one and build evals for each of them. It's not 100% coverage, but it is a priority for us to increase the coverage. Simon: How much interaction is there between the Claude Code team and the teams at Anthropic who are training the models in the first place? Is that quite a close collaboration? Cat: Across Anthropic, we all work quite closely together. We meet often to talk about what we expect the next generation of models to be able to do. Our research team has also been amazing about showing this publicly — we often talk in our blog posts about how we're targeting ever-increasing longer-horizon work , and how we train Claude itself to be honest, harmless, and helpful. We also put a lot of effort into making sure it's aligned with your intent, even if your intent is expressed in a fuzzy way. Of course, try your best to be specific about what you want, so Claude has all the context — but even when you're not specific, we teach Claude to make good assumptions. It's been a productive partnership. So many useful prompting tips in this section! Simon: Thariq, you mentioned this morning that the system prompt for Claude Code has been reduced by 80% because of Claude Fable . Can you go into a little more detail? What kind of things have you been able to drop? Thariq: It wasn't just Fable — it was Opus 4.8 as well, and going forward, future models. We have different system prompts for different models now. One of the patterns we saw is that we were over-constraining Claude. The initial, maybe Opus 4-ish models wanted a lot of examples, and removing examples was extremely helpful , because it was just more creative than the examples we gave it. Simon: That's really interesting, because one of the top prompting tips I give people is: give it examples. If that's no longer true, that kind of breaks my prompting model a little bit. Thariq: Same here — I was surprised to hear that. I think now it's more about the shape of what you give it — the tools you give to Claude, your system prompt, things like that. The other thing we did is try to give it more context and fewer "do not do this" instructions, because that's a very strong impulse for Claude, and especially if it conflicts with user instructions later on, that can be extremely confusing to Claude — "I've got this skill that says this and the system prompt says this." So we try to have fewer hard constraints, more context, and fewer instructions overall . It's definitely a science — it took a bunch of evals to build. Cat: In general, when you're prompting these models, you should always think: are there edge cases to the instruction that I'm giving it? When we went back and reviewed all the instructions in the Claude Code system prompt, we found a few cases where yes, this statement is 90% true, but there's a real 10% of cases where it's not true . We didn't want to constrain the model, or confuse it into thinking it should always do this. One good example is verification. Everyone here wants Claude to verify its work, and we had some instructions in the prompt that said: if you make a front-end change, always verify. But there's a limit to it. If it's changing copy from one string to another string, and the user says "just make a quick fix and update the test," maybe you don't want to verify. So we've adjusted our wording from "always verify, verify, verify" to something like: most of the time when you're doing front-end work you can't fully understand the experience by hitting the backend endpoints, so when you make larger changes to the user experience, please run the app locally. And in fact, that instruction probably isn't even good either, because what is a large change? Maybe it should test small changes too. In general, whenever you give a prompt to the model, you should think about the ways in which it could be misinterpreted by a well-intentioned human , in order to better understand how the model might interpret it — and soften the prompt so that it's actually 100% accurate, because you're giving this prompt to the model 100% of the time. Simon: What's fascinating about that is you're relying on the model's judgment — and that's got to be an Opus/Fable-level thing. Models a year ago did not have the level of judgment necessary to decide whether they were going to test a change or not. But that does break down if you're building for a wide range of models and trying to run the cheaper models for cheaper tasks. Cat: We actually have a different system prompt per model now , for this very reason. It's only our most frontier models that have this 80% token decrease — the older models still have the full system prompt. Simon: Do you think Fable and Opus are smart enough to prompt Haiku with more details, because they understand that Haiku has less judgment, less taste? Cat: We haven't been able to eval it — we don't have any hard data to show it. Thariq: There's a tough thing with smaller models sometimes, because sometimes the larger models can be more token-efficient on a hard problem than the smaller models . So there's a bit of intuition to build there — sometimes you really just want frontier intelligence almost all the time. The Pareto curve shifts, and it's hard to find. Simon: A year ago I did not trust a model to write a prompt. Today the good models are very good at prompting — a lot of my prompts are written by models, which feels absurd but works really well. What helped me come to terms with that was thinking about subagents, which are entirely about a Claude model setting up a prompt for another Claude model. Thariq: Workflows are actually a really good example of this, because it's Claude not just prompting a single subagent, but prompting the orchestration of many subagents, and each one of them gets a very detailed prompt. It's almost a level above just spawning a subagent. I've also been using it on my personal machine, giving it the Gemini API and saying: here, generate images . It's way less lazy than I am at prompting an image model. It's just Claude prompting Claude all the way down. Cat: I think Claude also wrote the prompt for the workflow tool . Simon: I've read that prompt — it's a good prompt. That's actually a frustration I have with Anthropic generally: you publish the prompts for Claude Chat , but you don't include the tool prompts and the Claude Code prompts. I still have to run a proxy to intercept them. I would love it if the Claude Code prompts were deliberately published — they're the documentation. They're how you know what the tool can do and how it works. Cat: I'll write down that feature request. I'll have Claude Tag do it. Interesting to note that OpenAI's prompting best practices for GPT-5.6 includes similar advice for their latest models: Favor leaner prompts Removing repeated instructions and examples and simplifying tool descriptions can improve task performance and token efficiency. In a sample of internal coding-agent eval runs, configurations with leaner system prompts improved evaluation scores by roughly 10–15% while reducing total tokens by 41–66% and cost by 33–67%. Simon: Claude Code is basically a big bag of tools. What's your bar for introducing a new tool? How do you decide when it's worth doing that additional engineering at that level? Cat: Do you want to take it? You introduced one of the best tools we have. Thariq: My career peaked when I introduced the ask user question tool. It's really hard. Especially for some tools — ask user question is Claude's tool to ask you — so it's hard to eval, and sometimes it's more of a user preference thing. Back then we had fewer evals, so it was very dogfooding based — or "ant fooding," our ant version of that. But overall we've been trying to trend towards fewer tools . The last set of tools we introduced was the task tool, I think — and we try to give Claude more general versions to do things. I have a long-running fascination with file editing tools - they were the subject of the old Aider code editing leaderboard , and I've watched with interest as they've evolved in different coding agents from search-and-replace based to line-number-based to more complicated patterns. The Claude API docs describe a text editing tool that's recommended for building against the API, but Claude Code seems to use slightly different approaches here. Simon: One of the most interesting tools is the file editing tool — you can have file editing as a tool, or you can tell it to use sed and grep and do things that way. What's the latest evolution of your file editing tool? Thariq: We still have one, but for example we removed our grep and other search tools — glob tools — in favor of native bash. Like I said in my talk earlier, the models are kind of more of a biology than a physics , and tool design especially is quite hard. I'm not sure if Cat disagrees and thinks there's a science to the eval of it, but I think tool design is more of an art, maybe — or a biology. Cat: I largely agree, but in general as we introduce more tools, we try to keep the cardinality pretty low and make sure that every tool we add has a distinct function from every other tool, so that Claude can very easily distinguish when to call each . For file edit, the reason we have it is actually because we can render it. We show people when Claude makes a file change, and there's this nice dedicated UI that says: do you approve this edit to this file? The reason we had a dedicated file edit tool was so that we could deterministically know that Claude was making a file change, so we could show people this nice UI. A lot of new users onboarding still really like this experience, so we've kept it around. But for a lot of us who are on auto mode right now — hopefully you're not on YOLO mode — I don't think it actually matters, and we could probably just remove file edit and be totally fine. It's the prompt injection question! Who better than Anthropic employees to explain how Anthropic sees the risk of prompt injection attacks causing their Claude Code instances to run amok? It turns out they really trust their auto mode - and see that as the feature that enabled Claude Tag. Simon: Let's talk about safety and security. I am deeply aware of the risks of prompt injection, and there are so many bad things that can happen if somebody else tells my Claude Code what to do. I still mostly run Claude Code in YOLO mode and feel incredibly guilty about it. What's the advice within Anthropic for safely running Claude Code? Cat: Why not auto mode? Simon: I am starting to use auto mode, but I don't understand it enough to get how safe it is. As of maybe three weeks ago, I'm defaulting to auto mode. Cat: Broadly within Anthropic, almost every single person uses auto mode. It is the best way to do long-running work in Claude Code while being safe. We've done extensive bashing. We have thousands of evals. We've commissioned many red teamers to create adversarial environments in order to trick Claude Code into doing bad actions, and we've mitigated every single issue that they found. We're going to publish some evals in the coming weeks, but we've pretty much mitigated every attack. Simon: That is a big claim. Cat: We'll share the evals for it so folks can assess, but we've been extremely diligent about identifying all the ways in which Claude might mess up and then updating auto mode to counter it. It doesn't catch 100% of things — that would be way too strong a claim. But for the main categories of risks that we're concerned about, like prompt injection and data exfiltration, the risks are far lower than the average human reviewer . I am very much looking forward to learning more about their evals and approach to verifying auto mode. Thariq: A little on how auto mode works — it's useful to build this mental model. Whenever Claude is doing a turn, or a bash call, there's a Sonnet classifier that is judging the tool call and also the context of the conversation — your instruction. There are some things around permissions that are dependent on your request: you don't want to give git push permissions all the time, but if you say "push this to GitHub," you want it to do it — and if you say "don't push," you want it to deny it. Auto mode will do that. That particular thing happens to me a lot, where Claude tried to do something because it's very helpful and proactive, and auto mode saw "don't do this" and surfaced it. So it's good at the dynamic permissions that you yourself give inside the prompt, which I think is really important. It also works well with our sandboxing infrastructure , because sandboxing is one of those things where there are so many different edge cases that it's hard for us to deterministically follow them. We have a sandbox, and when something needs to escape the sandbox — like a network request — auto mode can look at that request and ask: does this make sense? — and allow it. Simon: I hadn't realized auto mode is interacting with the networking sandbox as well. Cat: It interacts with any permission prompt the user would otherwise see. Simon: How old is auto mode? As a feature I had access to, it's only a couple of months old, right? (It was first made available to the public on March 24th .) Cat: We've been using it within Anthropic since January , so we've been hardening it for quite a while. Anthropic is extremely focused on safety and security, and we've been working broadly across our alignment and safeguards teams to enable the rollout internally, build out these evals, and make auto mode even more robust before sharing it with the world. Thariq: This is also the reason Claude Tag is so good — Claude Tag uses auto mode . I've heard a lot of build-versus-buy questions about a Slackbot, and I'm like: please, you probably shouldn't build your own AI Slackbot. There are so many attack vectors. You have a feedback channel that users can post feedback into, and now your bot is reading it. The work we've put in with auto mode — and we have a general Swiss cheese defense for security; we also RL against this stuff — I think this is really what makes Claude Tag work . It works seamlessly with your permissions, and you don't want to be prompt injected in your Slack. Simon: Are there any more security things in the pipeline that go beyond auto mode? Thariq: I think we're very secure. With Claude Tag you can provision your own credentials for Claude , so it doesn't need to act on your behalf — you can have Claude as an identity, and that also makes it easier to audit and inspect what Claude is doing. Simon: Because Claude Tag is influenced by anyone who can talk to it — it's got a much wider pool of people telling it what to do. Thariq: That's right. And of course we have probes as well with Fable, which is a downstream effect of our safety and research work. I think this is the moment where you see Anthropic being an AI safety company really paying off: we really want Claude to be able to run in an aligned way over long periods of time , and auto mode has to be basically flawless for this to work — it's all downstream of our being an AI safety company. Cat: We also launched trusted devices for the remote control users out there who want to be safer. And for all of our remote environments, we support credential injection . If you want Claude Code to be able to access Datadog, but you don't want Claude Code itself to hold the Datadog credential, you can set up our identity and credential management system so that the Datadog credentials are only usable by the agent but not accessible by the agent — we insert them on the fly when the agent tries to make a Datadog request. I really like that credential injection pattern, where Claude Code can access an API via a proxy and that proxy both audits the request and injects the relevant API key - so Claude can access authenticated endpoints without having access to the API credentials itself. Thariq talked about a sense of grief brought on by Fable-class models in his keynote in the morning, and we dived further into that as part of our conversation. I've been calling this Deep Blue . Simon: Let's talk a little bit about the human element. A lot of people are feeling a sense of loss now that so much of what they considered to be their role in building software is being subsumed by the models. How do you think about that? How has the past year and a half changed the way you think about your own craft and the value that you add? Thariq: Cat and Boris are such good reminders that you have to be more ambitious. They're always like: we're growing so fast, we have to be on the edge, we have to do the best work we can. That's a constant reminder for me — any time I'm slow on something, I'm like, okay, can I do it faster? Can I be more ambitious here? And oftentimes the answer is Claude, because Claude is getting better as you go — the last time I tried this, it was with the previous model. On your point about loss: I think this is real. If you're only trying to do the same work you were doing before LLMs, and now it's a prompt, it is, I think, kind of a sad feeling. And the way you offset that is by being more ambitious. I think Jared is such a good example — he hand-wrote all of the Zig code in his Oakland apartment in about a year, barely left his house, and had so much fun doing that. Now I see him rewrite all of Bun into Rust and he's having so much fun doing that — it's so much more ambitious, and that's how he offsets it. Generally it's asking how do I do the bigger thing and do more — I think success is fun . It's changing your ambition. "The way you offset that is by being more ambitious" neatly captures where I've landed on this issue myself as well. Simon: And Cat, what does that look like from a product management perspective? Cat: I feel like the product role just changes every single month. All the PMs on our team are this mix of engineer, designer, PM — most of them actually used to be full-time engineers. For us it really means plugging in whenever there's any kind of gap . If we have an idea and we didn't inspire any engineer to go build it, then we should just build it, put it into a notebook, and inspire people to take it to production. If the designs look a little off, let's take a page that's similar, do a first-pass design, and tag in someone who's very detail-oriented to fill in the gaps . Or if we notice that our team and product adoption is bigger within the company, and more people need to know what's coming down the pipe for Claude Code, Claude Tag, and Cowork — let's automate figuring out our whole launch calendar, let's automate getting those status updates asynchronously so we're not bugging people, and make sure our updates in our internal announce channels are fully detailed and to the point. For us it's very much understanding what the gap is right now between a great idea and getting something to our customers , and how do we automate it as much as possible . This reflects something I've noticed: when you can produce code so much faster, time spent blocked awaiting a decision from someone else becomes a much more notable bottleneck. Engineers who can make product decisions can move a whole lot faster, and the cost of getting one of those decisions wrong is much less prohibitive. Simon: What's a moment when Claude has surprised you? When the model did something you didn't think it would be able to do? Thariq: I've posted a lot about Claude video editing, but most recently I gave a talk at the ACM Agentic conference, and I asked, "Hey guys, do you have the edited video? I'd love to post it and share it with my comms team." They said, "Oh, it's taking so long." So I asked for the raw files. They sent me the video of me talking on stage, the video of the deck, and the audio file, and said, "Good luck." I gave this to Claude, along with my HTML deck, and said, " Hey, can you just edit this together? " And what it does is honestly incredible — I'm ready to ship it. It transcribes the entire video. It notices that sometimes the video of my deck is a little weird — there's a popup of an auto-update in the middle — and it goes, " Oh, I probably shouldn't use the video of your deck. What I'm going to do is slice it up, figure out which slide you're on, and use the HTML source instead. " So it displays the HTML source. Then it's got video of me, but I'm only taking up a small part of the stage, so it's cropping dynamically to where I am on the stage — and I'm pacing, so it's tracking me as I pace. And it's transcribing what I'm saying. Simon: This was Fable, right? Thariq: This was Fable, yeah. It was a good prompt, but it was a one-shot prompt. Then I asked it to add some interesting animations and graphics, and I was just blown away. It does ffmpeg, it does Remotion. Here's Thariq's video on how he used Fable to edit Fable's own launch video , and here's that launch video . I'm embarrased to admit that I've been finding it quite hard to come up with tasks that frontier models like Fable 5 and GPT-5.6 are unable to accomplish. Cat still doesn't rate its UX design skills: Simon: What can't it do? What are the things where you're still disappointed — where you're waiting for Claude Fable 6 to figure it out for you? Cat: I want it to have better design and UX taste. It's now at the point where if I write out a prompt with a detailed spec of how I want a feature to behave, it will usually behave that way. But the paddings might be off, or the interface just isn't delightful yet. It leans on existing best practices for how apps are designed, but for frontier AI products, there are so many new interaction experiences that we have yet to design . Simon: There's an Opus aesthetic — you can look at something and go, "Yeah, that was designed by Opus." It'd be good if we could move beyond that. Cat: Yeah. I'm very excited for future models to hopefully be interaction design thought partners . Thariq: What can't it do? I would love to see it interact more with the real world. Can it solve science? Can it orchestrate the experiments? There's some amount of coding that goes into that, but there's also this other taste of the broader world that it needs. I figured this would make a great closing question: Simon: Which parts of Anthropic's company culture do you think uniquely help Anthropic be productive with these tools, that other companies should steal? What are the cultural hacks people should be adopting from you? Cat: I'll share one for Claude Tag. Claude Tag works best when you have it in a public channel, and when most of your channels are public. Claude Tag is able to search across all public channels to get as much context as possible to give you the highest-accuracy answer — and it's only able to do this if it has access to everything . Thariq: I mentioned this in my keynote, but it's so important to me I want to re-emphasize it. The co-founders say we don't negotiate against ourselves , and I think this is really important. You can imagine trade-offs in your head and talk yourself out of doing something ambitious — or you can just try to do the ambitious thing. We're so often asking: what if we just did it? Is this a real trade-off or not? And if so, why — where's the proof that it's a real trade-off, and not just something that sounds reasonable? Make the trade-offs show themselves to you. Be as ambitious as you can. I couldn't resist throwing in this one as well. Simon: What's one of your favorite absurd things that you've built with Claude, just because you could build it? Thariq: I'm working on a 2D Street Fighter fighting game with me as a character — and my friends as well. It uses Claude Code to prompt Gemini — and honestly the Seedance model is pretty good — to make video animations. It works great; it's so good at prompting, and it can verify the frames to check whether an animation was good. Simon: Is this Street Fighter 2-level 2D sprites you're generating? Thariq: Yeah, exactly — 2D sprites. The animation looks amazing. And it can also figure out hitboxes — it can be like, "Oh, your fist is here, I'll draw the JSON hitbox." It's incredible. Cat: Mine is much more simple. I'm a big rock climber and a lot of my friends climb, so we have this little app we built with Claude Code where we log all the projects we're working on. We also go outdoors together a lot, so we have Claude do all this research with workflows. Workflows is amazing — we brand it as a coding tool, but it's amazing for doing deep research for travel. I also plan our team offsites, and it's good at finding venues that can fit all of us. I use workflows to research all the climbing destinations we might want to go to, and what has direct flights from where all of us are located. It goes to Mountain Project and finds all the climbs at our grade level. It finds the Airbnb. And I don't like hiking, so I care a lot about it having a very short approach — very short walking distance from where the car parks to where the rock actually is — and it filters for this. With existing apps I have to manually click through Mountain Project, but with this I just put in all of our preferences and it's a custom app for us. Simon: So you're basically vibe coding Jira for mountain climbing. Cat: Exactly. We had a few minutes at the end for questions from the audience. Audience: Do you have any near-term plans to build more eval tools for us to build eval datasets, and more observability tools to monitor the performance of agents and workflows? Cat: We've considered building eval tools, but I think the limiting factor actually tends to be that it takes a long time for customers to build really high-quality evals . So I think the tooling is less of the constraint, and more the skill set of how you build a great eval. That's an area where we're excited to both invest internally and hopefully share some best practices externally. Audience (Sai): I'm interested in the memory and the multiplayer. How is memory being designed today? I assume it's around files. And second, have you thought about an orthogonal direction where you would actually need a data store for these memories, instead of files, to scale it better? Thariq: Right now for Claude Tag the memory is channel-specific. Every Claude in that channel has a shared memory, and the instances have a session — but the session can contribute back to main memory. We do a lot of memory research, and it can be kind of unintuitive what the right way to do memory is. We're always running memory experiments. How it works right now in Claude Tag is a markdown file per channel. You are only seeing the long-form articles from my blog. Subscribe to /atom/everything/ to get all of my posts, or take a look at my other subscription options . Claude Tag (Claude's new collaborative Slack integration) now lands 65% of the product engineering PRs for the Claude Code team. Claude Code ships features to Anthropic employees first, and only ships the features that demonstrate user retention with that cohort Critical changes to Claude Code are still reviewed manually, but the team increasingly relies on automated code review for the "outer layers" of the product. Adding examples to a system prompt is no longer best practice for models like Fable 5 or even Opus 4.8. The Claude Code system prompt recently reduced in size by 80% . Likewise, lists of " don't do X and don't do Y " can reduce the quality of results from the latest models. Dogfooding inside Anthropic is called " ant fooding ". Anthropic really believe in their auto mode , and see that as an enabling technology for Claude Tag. Thariq advises offsetting coding-agent-induced Deep Blue by " being more ambitious " with the work you take on. Fable is competent at editing video , and Thariq used it to edit its own launch video. Anthropic's culture of working (internally) in public is key to their success, as demonstrated by the way they use Claude Tag in their public Slack Channels.

0 views
Julia Evans 2 weeks ago

Some more things about Django I've been enjoying

Hello! I’m on a funny journey right now where I’m trying to learn how to make websites in a sort of 2010 style, where I have an SQL database and render some HTML on the backend. It’s kind of an interesting journey because it doesn’t necessarily feel “easy” to me to make websites in this way: I never learned how to do it in the 2000s or 2010s, and there’s a lot I need to learn. So here are some Django features that make building this kind of site feel more achievable than when I was trying and failing to use Go’s standard library or Flask. And I’ll talk about a couple of issues with Django I’ve run into. Previously the toolkit I felt confident with for making websites was: I really liked this frontend-heavy approach for these super simple applications but when I started thinking about making something with a lot of different pages (instead of literally just one page), I didn’t feel so excited about the options I saw that involved a lot of frontend code. So I figured I’d try the backend. Writing a backend-focused site that uses as little JS as possible feels the same to me in a way as writing a single-page JS website that does as little on the backend as possible, even though they might seem like opposites. In both cases I’m just trying to keep as much of the logic as possible in one place. Now for some thoughts about Django! I learned that I can define a “query set” class in Django with a bunch of methods with different statements I might want to use while constructing a query: Here’s how I use it in my view code once I’ve defined what all the methods mean: and here’s how I define the methods: The syntax for defining the filters isn’t my favourite, but I spend most of my time just using the methods, and it feels super readable and nice to use, and it makes me want to look into other query builder libraries in the future. In the past I thought “I know SQL, who needs a query builder?”, but this kind of structure does make it really nice to read. I found an example of someone who wrote their own small query builder in Python that I want to read later to think about whether I would enjoy using a more minimal version of this. There are a bunch of little quality of life filters available in Django templates that are super useful for generating HTML. The ones I’ve used so far are: These are all small things individually but I feel like it makes a big difference somehow to just have them available. I think my favourite template filter is : in this site sometimes we use filters like to decide what’s displayed. that will make a link to the same query string with one change, like this to link to the previous date: Or to remove the parameter: I still really love Django’s automatic database system. It’s amazing to be able to just edit a model to add a new field or whatever, and then Django automatically generates the migration. So far we have done 19 database migrations and I think there will probably be more! It makes a huge difference for me to be able to just easily change the database as my understanding of the problem changes. Django’s documentation sometimes offers the option of using class-based views and inheritance to organize the code in your views. For example I have four views that share a lot of code, and I could use inheritance to manage that by defining some kind of parent class and then having my other views inherit from it. I tried it out and I did not enjoy the experience of using inheritance to share code between views. I switched to using functions instead, sort of how this post advocates, and that was a lot more straightforward. I’ve never had a good experience using inheritance in Python and I don’t think I’ll try to use it again. But I don’t mind using inheritance to use the interfaces Django itself provides: for example if I want to define a query set I need to write something like . I don’t think too hard about it and it seems to work. (as a meta comment: I’ve been working on talking about my programming opinions by just saying “THING does not feel good to me, I prefer OTHER THING instead”. That post I linked to says that function-based views are the “right way”. I’m not very invested in whether it’s “right”, but it’s validating to know that other people feel similarly to me about inheritance) At some point the LLM scrapers discovered our site, and started sending us maybe 10 requests per second. I blocked them which is working for now, but it made me think about what the site’s capacity is. I’m used to writing Go backends where the performance situation is pretty straightforward (usually everything is just fast enough), and a Django site is very different. Some light load testing (with ( ) shows that right now we can serve about 2-3 requests per second (on a ~$10/month VM). It’s tempting for me to go down a rabbit hole where I do a bunch of profiling to figure out what’s slow and try to make it faster (there’s py-spy for that, and py-spy is great and super easy to use, and profiling is fun!) But I really don’t understand what I should expect in terms of performance from a Django site and how I should be thinking about at a higher level. Some things I haven’t figured out yet: I think one thing I’m learning about Django is that because it’s a Framework (tm), it’s easy to accidentally misconfigure it. For example, when I was thinking about why my site was slow just now, I read the django performance docs and I noticed a comment saying: Enabling the cached template loader often improves performance drastically, as it avoids compiling each template every time it needs to be rendered. When I’d done CPU profiling I’d noticed that it was spending a lot of time rendering templates! Maybe this could help me! Clicking through the link, I saw that the cached template loader was supposed to be on by default, but I’d turned it off by accident while trying to do something else. I think this “I turned off the cached template loader by default” things is an example of how I still find the django settings file to be pretty confusing and difficult. I guess I should just be careful when I go in there. After turning on template caching, it seems like the site can now pretty easily handle 12 requests per second or so without using all of the CPU. I have not carefully benchmarked the before and after but it seems like it’s made a pretty big difference. One thing that’s been surprising to me about Django performance is that I’ve always heard the advice “if you have a performance problem, check your database queries! Maybe add an index!”. But I’ve been running into a variety of performance issues (like this template caching thing) that are not because of slow queries, so instead it’s been more useful for me so far to start by running a CPU profile. And since I’m using SQLite, any slow database query problem will show up on the CPU profile anyway. Anyway I don’t want to get too far into site performance. Like I said it’s easy for me to get interested in profiling, but actually I know a lot about profiling and it’s not the most important thing for me to learn about. I might say more about what I’m enjoying (or having a hard time with!) about Django later. Trying to write some shorter blog posts recently. static site generators (like for this blog) static sites that do some fun stuff with Javascript (like this sql playground ) simple Vue.js single page apps with either a Lambda as a backend or a Go backend (like mess with dns ) translating plain text URLs into links, or line breaks into ( ) formatting dates ( ) , which takes a Python dictionary and automatically converts it to JSON and inserts it into the HTML as a tag in a safe way If I have a site that’s going to be getting occasional bursts of traffic, do I want to be able to scale up? Do I want to design the site so that more things can be cached? (and do I really have to? caches are so annoying to get right!) The django performance docs say that Jinja is faster for templating, do I want to think about switching templating systems? Those docs also say “{% block %} is faster than using {% include %}”, I wonder if it’s a big difference and if so why

0 views
matklad 3 weeks ago

Memory Safety's Hardest Problem

Uplifting a lobsters comment for easier reference. The central memory safety counter example, the hardest case to solve, doesn’t have anything to do with destructors or heap: This sort of example also breaks Ada: https://www.enyo.de/fw/notes/ada-type-safety.html We have a tagged union, which can hold either or . We initialize the union as , take a pointer to its internals, overwrite the original with , and then use the pointer. The pointer is still typed as , but the bytes it points to now belong to : a type confusion. This being said, we care about memory unsafety primarily because it leads to exploitable software, and it’s unclear just how impactful the example above is in practice. It is a happy coincidence that by far the most exploitable memory error in practice, the infamous buffer overflow, is also trivial to fix with compiler-inserted bounds checks. The biggest miss of the industry when it comes to memory safety is not listening to Walter Bright: https://digitalmars.com/articles/C-biggest-mistake.html I bet that, had we got syntax around C11, quite a few issues wouldn’t have happened! See also What is Memory Safety?

0 views
Maurycy 3 weeks ago

Regressive JPEGs:

One of the cool features of JPEG files is that there's the option to save low frequency components first. This means that a partially downloaded image will be displayed at low resolution instead of being cut off. In the file, this works by breaking up the compressed data into multiple "scans", each prefixed with a header. Here's the first scan of a representive image: ... this one includes the lowest (DC) Fourier bin for all three color channels. The three color channels are YCbCr instead of the usual RGB. The luminance (Y) seperated because it must be high quality, but the color can be fudged quite a bit while looking fine. Very roughly: Y = G, Cb = B - G, Cr = R - G After it, the file contains eight more scans to fill in the rest of the data: Scan number Channels DCT bin range Precision 0 Y Cb Cr 0 - 0 Half (-1 bit) 1 Y 1 - 5 Quarter (-2 bits) 2 Cb 1 - 63 Half 3 Cr 1 - 63 Half 4 Y 6 - 63 Quarter 5 Y 1 - 63 Half 6 Y Cr Cb 0 - 0 Full 7 Cr 1 - 63 Full 8 Cb 1 - 63 Full 9 Y 1 - 63 Full Scan #0 contains a very low resolution preview of the image. Scan #1 adds some details to the luminance. Scans number two through five contain full low precision data. Scan 4 has an unusual spectral range because it's filling in the gap left by #1. That way, number 5 has full quarter precision data to build on. Scans six through nine add the final missing bit to bring the image to full quality. Given what I said about color being less important, it might seem weird that my example has the color data first: This works because the the chrominance is saved at half resolution (quarter pixel count). As a result, full chrominance data (Cr + Cb) only weighs half as much as luminance. Since each scan explicitly sets its spectral range , it should be possible to construct a JPEG file where future scans overwrite already rendered image data. Actually, it's very easy to do this: Concatenate multiple images with the same resolution and filter out the start-of-image, start-of-frame and end-of-image markers. This can be done in a hex editor, but I used a quick and dirty C program. When served over a slow network , this concatenated file will switch between multiple images: Click to open in new tab But, most decoders will give up after some number of scans : I think this is done to avoid a zip bomb style problem... but it prevents this from working on more than 9 frames, which is not enough for a proper animation. To do that, I'd have to minimize the number of scans in each frame. The simplest idea is to start with baseline JPEGs that only have a single scan. ... but it doesn't work: In progressive mode, a scan can't contain both AC (bins above 0) and DC (bin 0) data at the same time. This limitation doesn't exist for baseline mode, but the baseline decoder stops after the first scan. Since AC data must follow DC data, the smallest possible "progressive" JPEG contains a single DC-only scan. Because the DCT runs on 16x16 blocks, such an image won't a solid color: it'll be 1/16th of the original resolution. Scan number Channels DCT bins Precision 0 Y Cb Cr 0 - 0 Full Doing this, I can get Chrome to render around 90 frames before giving up. Other browsers like Firefox have more patience, but a 90 scan image seems to work almost everywhere. As a bonus, this avoids the ghosting of the naive attempt: that happened because AC scans are supposed to refine old data. Normally, this allows images to include multiple precision levels without inflating file size... but doesn't play nicely with my tricks. If the file only includes DC scans with no actual progression, this isn't a problem. Since a "DC-only" frame is a standards-compliant images , creating them doesn't require anything special: Using these, it's possible to pack a whole video inside a single image: Click to open in new tab Besides unconventional rickrolls and other trolling, this has no practical applications: there's no way to add timing information, so playback is entirely dependent on network delay. ... although there is a lot of fun to be had using partial rendering: This is a pure HTML video using <dialog> tags: badapple.rose.systems Of course, there's no rule that the data must be hardcoded: here's a interactive single-page application with no CSS or JavaScript. (seems slighty broken, I'll investigate later) Related : /projects/bad_jpeg/merge.c : The code used to generate these images /projects/bad_jpeg/merge.c : The code used to generate these images

0 views
James Stanley 3 weeks ago

Optimistic epsilon-greedy

I've been working on optimising revenue on my Countdown website the last few days. I have had a Countdown solver tool online since about 2009. It is to this day the most popular website I have ever made, it currently gets about 70,000 pageviews per month. The site has been earning revenue from AdSense for years. Up until last week the site was just 2 static HTML pages: one for the Countdown Solver and one for the Countdown Practice game. That didn't give me much opportunity to run experiments on the site, and I never really had the inclination to try. It was basically a web program . But now, LLMs to the rescue. I now have a Python Flask application serving the site, and a lot more related information pages for people to read. And serving the site with an actual web application means I can run experiments like A/B tests to see if there are changes I can make to the site that cause people to stick around longer. And therefore look at more ads. A good alternative to A/B tests is multi-armed bandits . Instead of splitting your traffic equally between the different variants you want to try out, and then waiting to collect data, and then picking a winner, you have the site automatically determine the winner on a continuous basis, and show the winner 90% of the time (greedy), and a random selection the rest of the time (epsilon). I am using a multi-armed bandit to decide which "info" pages to suggest at the bottom of each page, and also to decide which Amazon Affiliate links to show. (Yes this is all very grubby, what can you do?). The "winner" is the choice that has the highest click-through rate. So for each choice we need to track how many times we've displayed it, and how many times it's been clicked on. If your reward function is more complicated you might find it more complicated. Steve Hanov's blog post on multi-armed bandits, linked above, goes over the case where you might worry that a particular variant gets a click early, just by random chance, which gives it an apparent high click-through rate, which then means the site is going to show that variant to everyone. And that's not actually a big deal, because showing the apparently-high-performing variant to 90% of traffic gives it a lot of opportunities to prove that it's not actually that good, and it's click-through rate will come back down. A much bigger issue, in my opinion, is when you add a new variant. Let's say you already have 9 variants that all have click-through rates around 1% and have had about 1000 views each. Then you add a new variant. This new variant starts out with 0 clicks. Now you have 10 variants, 9 of which have a CTR of 1% and your new one has a CTR of 0% (technically a degenerate case with 0 views, but becomes firmly 0% after the first view). And let's say your site expects 1000 views per day. 90% of the time your site is going to be showing one of the old variants, because no matter what happens to their CTRs, they can't go below 0% , so they will forever look better than your new variant. The remaining 10% of the time your site is going to be picking at random amongst all variants. So your new variant is going to get about 1% of your traffic. Or 10 views per day. If your new variant also has a CTR of about 1%, then you'll expect to get about 1 click per 100 views. If it is only getting 10 views per day then it could easily be 10 days before you get the first click, during which time you're not even gathering much data on it. So what I'm doing instead is defining the CTR to be (clicks+1)/(views+1) . That is, we always optimistically assume that the next view is going to get a click. That means a new variant starts out with a CTR of 1/1 = 100% . We skew the selection towards those variants that have not had many opportunities to prove themselves yet. In this case the new variant will get 91% of the traffic until its optimistic-CTR falls below that of the next best variant. That could easily happen within the first day of releasing the new variant, so this "optimistic epsilon-greedy" algorithm broadly behaves exactly the same once the number of views is high enough, but it discovers the true CTR for newly-added variants much more quickly than the standard algorithm. Even if the new variant actually never generates any clicks, its CTR drops below 1% within about 100 views ( 1/101 ) so it won't be taking much traffic away from your older variants if it doesn't work very well.

0 views
seated.ro 4 weeks ago

You fail to learn if you don't learn to fail

If all your time is spent watching output tokens, where do your input tokens come from? Letting an agent rip on full auto is basically doom scrolling. Even worse if you're doom scrolling while the agent runs. We humans love frying our dopamine receptors. This feels great until you realize what you were offloading: the struggle. The part where you fail. Failure is the entire point. You don't make progress in the gym unless you take a set at least close to failure. The muscle only adapts when it's forced to. It is no different for the brain. It is very hard to admit to yourself that your skills have atrophied. It is even harder to admit this to other people. I will admit that over the past several months my brain has gotten smoother (and I wasn't even on Twitter much!). Recently, I had written an abstraction for my diff viewer ( diffy ), an element system with a macro that lets agents write html-like code in rust for native ui (they reason better with this). But it wasn't adopted everywhere in the repo yet, so when I asked for a new feature, the model decided to hand paint it straight to the viewport instead. Every behavior the element system gives you for free was just... missing. Text wasn't selectable. Hover highlights wouldn't go away. And since I wasn't looking closely, it iterated on the slop and produced more slop, more bugs. I just kept saying continue. I lost a whole day untangling it, and the funny part is that once I actually looked at what it had built, every bug was the same bug. When you hit a roadblock and your immediate reaction is to reach for something else (previously, this used to be other people, but now it is a language model) you are essentially skipping the part where you actually learn to solve the problem. It is funny how one of the best "learning tools" has turned out to be the number one cause (anecdotal. sue me) of the lack of learning! It's been a few months since I started writing this, and things have gotten more dire. Several major software services barely work now, grown engineers I once respected are writing somber posts about missing a language model that was banned for a while. Mourning. For model weights. It's all so dystopian. As the agents get better, one is basically expected to produce code at an alarming rate. The timeline to get something done is compressed but the time it takes to come up with solutions to hard problems has not. There are usually a few good abstractions one can come up with that balance the upsides and tradeoffs for most software problems. However it is currently trivial to turn your brain off and let the slop flow. The code will be complex. It might look like it all works, but something always breaks. And the solution to that? More slop. Software quality is collapsing as a result, and the societal expectation that engineers understand what they ship is disappearing. You never understood the code in the first place. So when you need to change it, you're asking the same stateless clanker to modify code it has no memory of writing. All output tokens and zero thinking tokens. A lower barrier of entry to write software doesn't imply the standards for good software must be lowered. The growing trend is to do things because you now can (supposedly), but we used to try and do things because we could not out of sheer stubbornness. Carmack and gang shipped QuakeWorld with client-side prediction over dial-up when the conventional wisdom was that twitch shooters over the internet were unplayable. This only happened because Quake's original netcode was laggy and everyone hated it. (They fixed it in a month.) George Dantzig arrived late to class, mistook two "unsolvable" statistics problems for homework, and solved them. Nobody told him they were impossible, so he just did the work. Andrew Wiles spent seven years alone in his attic working on Fermat's Last Theorem, a problem mathematicians had given up on for 350 years. He announced the proof, a reviewer found a hole in it, and he spent another year fixing that too. Notice that all three of them became who they are because of the struggle, not despite it. The people benefiting most from generative tools today, say Terence Tao or Mitchell Hashimoto, already put in the time, so when they offload work they're just skipping the typing. When people like you and me (if this is not you, then I apologize) offload, we skip the grind itself. With language models, easy tasks got easier, hard tasks stayed hard. The hard part was never the task itself. I don't know, I am figuring this out as I go. The amount of time I have spent actually programming has been dropping month over month this year. I used to have a coding stats section on my website that would track hours I spent writing code split by language, recently I had updated it to this: and it made me quite sad. I do think that sometimes all you need is to realize that the thing you are doing is actually detrimental to your growth. Consistency matters more than one would assume. If you consistently take some time away from these tools and actually use your brain, that alone is already significantly better than offloading your thoughts. Solve the problems yourself. Or at least try, fail, and spend time thinking. There is seemingly no "learning" phase anymore. You are expected to just know things. Learning is fun, don't let anyone take this away from you. I've written about this before . It is probably going to be slow, learning takes time and effort. You will feel stupid (I feel stupid). This is a good feeling, because there exists a world where you are no longer stupid and the path towards it is learning. Books still exist! Libraries are still open, notebooks waiting to be written in. Read more. Write more. If you really do care about improving yourself, be honest and use these models for what they are, highly efficient filters of zettabytes of data (the internet is estimated to be 175-240 zettabytes ( 10^{21} bytes)). It was extremely difficult to identify what one needed to read to learn niche topics even like 2 years ago. I remember asking a good friend of mine to recommend material to dive deep into learning about SIMD, and honestly there wasn't much stuff to read except the Intel Intrinsics Guide. And if you've ever taken a look at that, it is quite cancerous for a first-time reader. Language models are super useful here because you can point them at material and you can ask questions that pertain to the thing you care about and it will simply just tell you the correct things. One good thing in this age of slop is to consume knowledge at an unbelievable pace. I don't necessarily mean using only model output for learning (I don't trust them to learn any topic more than a shallow amount), but rather using them to help sift through the plethora of information available out there and identifying the right things to read. Human slop exists too and using a language model to supplement your learning might help keep you sane (ironically). I like using these models to write code that I tell it to write (outside of work I enjoy doing it myself entirely), and I am largely disinterested in asking it what I should write. There are exceptions of course, because not everyone is working on scaling software services which has largely been solved (but slowly being forgotten), but that would be for you to decide. The best model you have access to (and it has solved continual learning) is, and always has been, the one inside your skull. It's time to scale up its input tokens.

0 views

Establishing an Identity

If you’ve followed me on RSS for any amount of time, first off, thank you so much! Second, you may not have noticed how often this site changes. RSS protects you from the near-monthly changes that my mad scientist side makes to this site. This year alone, ThatAlexGuy.dev has been powered by 11ty, Hugo, plain HTML, Bear, Micro.blog , and Pure Blog. My files have sat on OpenBSD Amsterdam, DigitalOcean, and a Laravel Forge VPS. I’ve written new articles and lost old articles in migrations. My site has switched appearance more frequently than a Bian Lian (变脸) performer! I’ve come to realize I’ve been seeking both an identity and a voice. I want an outlet that reflects my interests, my background, and my day-to-day, but that’s more than what I could accomplish on something like Mastodon. All that brings us here, iteration 4 (or 8, or 15, or 16, I can’t remember). There are a few key differences and intentional choices that reflect where I want ThatAlexGuy to go. Building a new experience that will stick and satisfy the goals in my head won’t be easy, but here are the guiding pillars that are to shape what’s coming next. I have a desire to create in-depth, well-researched, and potentially interactive content. Many of my current posts come with a “1-minute read” tag. I want to change that. I’ll be digging into topics with greater detail, cross-referencing multiple sources, and (hopefully) interviewing others. As a result, I’ll be posting less frequently, but my new goal is quality over quantity. Regulars on my site will be aware of my “Photo Journal” series in which I posted a set of photos around a theme (macro, nature, Gameboy Camera ). I want to continue building my photography skills through the incorporation of high-quality photos in my articles. While text sets the tone, visuals set the atmosphere in an article. Here’s the big tomato, as they say (nobody says that): defining what this site represents. That means setting the tone and defining how topics string together to form a consistent narrative. I’ll be figuring this out for a while, but I want to leverage my interests such as indie technology, vintage computing, time away from the screen, photography, and Chinese culture. So what’s changed so far? Quite a bit! First, ThatAlexGuy.dev is now run by Ghost.org . For myself, this means less time in the technical weeds and more focus on writing. For readers, it opens the doors to a wider audience. Email newsletters are a more accessible way to stay up-to-date on new articles. Don’t worry though, RSS isn’t going anywhere! In fact, I managed to fix the broken RSS feed URLs from previous migrations (hopefully)! I’ve started to define the personality of the new site. I pulled background and accent colors from one of my favorite atmospheres in a game (Sprout Tower in Pokémon Gold). Using my iPad, I’ll be creating article images that give a calligraphy + hand-painted vibe. I’ve also brought in my Chinese name for the logo(小艾 - Little Alex). I’m working on my first longer-form article. It probably won’t be great, but first attempts never are. From there, I hope to refine my writing, researching, and supporting photography.

0 views
Unsung 1 months ago

“…or I could click seventy buttons.”

I like Angela Collier’s videos about physics and I was delighted to discover this 18-minute one … = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/or-i-could-click-seventy-buttons/yt1-play.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/or-i-could-click-seventy-buttons/yt1-play.1600w.avif" type="image/avif"> …because it’s a great continuation to the thread about the complexity of Microsoft Office I shared recently. Collier talks about why physicists prefer LaTeX to Word. LaTeX is sort of a nerdy HTML that predates HTML. It looks like this… = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/or-i-could-click-seventy-buttons/1.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/or-i-could-click-seventy-buttons/1.1600w.avif" type="image/avif"> = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/or-i-could-click-seventy-buttons/2.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/or-i-could-click-seventy-buttons/2.1600w.avif" type="image/avif"> …and given how nerdy HTML already is, you might imagine this is a power-user tool that’s chiefly about power and control. But Collier makes the argument that there are some things that LaTeX makes much easier: This is really interesting because it goes right to the core of the uncomfortable truth: naïve design decisions meant to make things easier might achieve the opposite. I shared the ForkLift example where the team didn’t understand what made the previous version great , and more recently the animation that could slow people down . (Of course, there is also the issue of typographical craft of LaTeX documents set in Computer Modern , but let’s save this for another time.) Also, the video starts with Collier apologizing for potentially making the audience feel dumb in a prior video. I don’t think it’s a joke, and I found it thoughtful and refreshing. #attention #complexity #enshittification #flow #youtube there is absolutely no need (or peer pressure) to spend time styling the document by choosing fonts, colors, etc., there is no “live preview,” and making a PDF is a separate step similar to compilation in coding – which means it doesn’t constantly occupy your mind, GUIs can slow you down because the keyboard is faster than the mouse, LaTeX doesn’t give you a lot of control over positioning, which is better than giving you only a semblance of control over positioning ( this is the TikTok meme Collier alluded to briefly ).

0 views
David Bushell 1 months ago

Astro is fine I guess

When I’m not fighting WordPress I deliver static HTML or the occasional JavaScript framework integration. For personal projects I have ‘fun’ with my own static site generator . This week was a side quest (soon to be main quest) to build my new company website. We’re talking proper business here so I can’t be messing about. I figured an off the shelf SSG would be most suitable. I asked the socials, “ 11ty or Astro ?” Both are popular but Astro had the edge. I gave Astro an early spin back in 2022 and found it slow . Maybe it’s good now? I ran with minimum release age to avoid immediately getting pwned . I selected Astro’s “Use minimal (empty) template” option and it generated both an and file — are you f — deep breaths, don’t fall for the rage bait. I code in a modern editor so I installed the recommended Astro extension. At first I struggled with Zed recognising HTML. I discovered a restart temporarily fixed the issue, but I guess I restarted one time too many because now the Astro LSP is completely broken. No modern comforts for me then. At least I can look at HTML without the red squigglies. I know what you’re going to say, “Dave bro, you’re inflicting this pain upon yourself! Just write HTML!” And I should. I just want native no-framework HTML includes , you know? Can you imagine the civilisation we’d live in if that could happen? I persevered and got my templates built with minimal fuss. I added a markdown collection and got the blog part blogging. It’s obvious that people use Astro to build real websites because all my “how do I” questions had an answer in the documentation. I’ve been forced to deploy way too many “React spaces” in my templates because Astro’s whitespace treatment is a mystery. I don’t need many components so I haven’t gone deep on Astro vs JSX . My site has zero JavaScript on the front-end. I plan to keep it that way. Edit: Christian Niklas on Mastodon shared a link to a recent Astro update where they added a option that defaults to no longer “following HTML rules.” Umm… okay. Set this to or if you’re building a website? I set it to . Minifying whitespace is over-optimisation. Astro has got the job done, despite the developer experience being broken out of the box. I dread to think what graveyard of dotfiles is installed if I choose a non-minimal start. I can easily de-Astro my templates should I need to. Right now Astro is solving the right problems and the issues are but a nuisance. Final conclusion: Astro is fine I guess. I’m not convinced Cloudflare’s acquisition is a good thing, considering their record for performative slop. I’ve lost my enthusiasm for DX and tooling to be honest. Even my own SSG experiments are collecting dust. I’d call the ecosystem a lost cause if I was being dramatic. I just try to avoid the worst of it and care about the end product: shipping a damn fine website! Which I can’t do because I’ve got more businessing to business before this particular site sets sail. Maybe in a few months? It’s looking awesome on though. Thanks for reading! Follow me on Mastodon and Bluesky . Subscribe to my Blog and Notes or Combined feeds.

0 views
Alex White's Blog 1 months ago

Go have fun with the web

Back in the days of Geocities, I spent a lot of time hacking away on raw HTML and CSS. I enjoyed tweaking things, making it just right and experimenting with random ideas I had. I’d sketch things out, then turn them into a close(ish) version on the web. “Under construction” gifs would hide my unlinked, mad scientist HTML files. As I grew older, the idea of “hustle” culture slowly killed out this mindset. Instead of having fun, I felt everything I do on the web had to serve a purpose. If I wasn’t building something that might make money, I was wasting my time. And guess what? In 15ish years of operating under that mindset, I’ve made maybe $500 online. Pretty terrible investment if you ask me. I’m willing to bet I’m not alone in this mindset, it seems embedded into the millennial DNA. We’ve grown up with stories of dot com entrepreneurs making it big while sipping Mojitos on the beaches of Chiang Mai. You’re always just a few more late nights from quitting your job, joining NomadsList and traveling the world! The truth is, you’d probably have a better chance winning the lottery, so why waste your time chasing the impossible? Why turn an artistic, creative outlet into a second job that doesn’t put food on the table? Embrace the web as a hobby. Like pencils, paintbrushes and clay, the web is a way to give “physical” form to the images in your head with HTML, CSS and JavaScript. When you stop building for scale, potential customers and imagined profit, you free yourself to have fun. Build silly, build simple and above all else, build for the sake of creativity.

0 views

Let AI Burn

If you liked this piece, you should subscribe to my premium newsletter. It’s $70 a year, or $7 a month, and in return you get a weekly newsletter that’s usually anywhere from 5,000 to 18,000 words, including vast, detailed analyses of NVIDIA , Anthropic and OpenAI’s finances , and the AI bubble writ large (updated to version 3.0 a few weeks ago). My Hater's Guides To the SaaSpocalypse , Private Credit and Private Equity are essential to understanding our current financial system, and my guide to how OpenAI Kills Oracle pairs nicely with my Hater's Guide To Oracle . This week, I published the Hater’s Guide to Softbank — a sordid tale of tech’s most degenerate gambler, who, thanks to a couple of early lucky wins, has managed to set the foundations for the AI bubble’s biggest (and possibly most gratifying) downfall. And, on Friday, I’m going to take a deep dive into the memory industry — and the reason why you can’t afford a new gaming PC.  Subscribing to premium is both great value and makes it possible to write these large, deeply-researched free pieces every week. Soundtrack: Mastodon — Streambreather No bailouts, no handouts, no special treatment, no tax breaks, no CHIPS act, and no sovereign wealth fund. It is time to tell the AI industry to go fuck itself, because it’s effectively done the same to the rest of society. This industry is unworthy — a sham conjured up by a tech industry that’s run out of ideas, a trillion-dollars’ worth of manufactured consent and entirely-avoidable financial crises — and should not be protected under any circumstance.  Every single time you hear somebody discuss “bailout” or “too big to fail” or “sovereign wealth funds,” know that this is the industry, on some level, attempting to create the air that it cannot die , when in fact every one of these companies is just as weak and brittle as any other startup. I also think that the media — and the world at large — is too ready to accept the prospect of a bailout after watching those who drove the world into a ditch in 2008 escape blame, and I must be clear: the AI industry is very different to the financial industry. It is inessential to the economy, and its relevance is only as large as the hype campaign that sits behind it.  This is an industry of losers that has inflated only because of the joint manufactured consent of Silicon Valley, the mainstream media, and an enshittified stock market that rewards grifting and circular financing . OpenAI had $5.7 billion and Anthropic a little under $5 billion in the first quarter of this year — and those revenues mostly came from companies that were burning AI tokens at a horrendous rate because they’d just been forced to pay the actual cost of AI — and now everybody’s pulling back on that spend .  Generative AI will not bring us AGI, nor does it do much of what we associate with artificial intelligence. It is not autonomous. It is not “intelligent.” It does not have thoughts, or “knowledge,” and no matter how many layers of harnesses and scripts you put on top of it, it is still ( per OpenAI ) mathematically certain to hallucinate. I estimate that at least 70% of the entire AI industry’s revenues are made up of OpenAI and Anthropic’s compute spend , and as both companies are horrendously unprofitable, this means that the AI industry is, for the most part, venture capitalists funnelling money to hyperscalers so that they can funnel that money to NVIDIA or data center capex. If this software were worthy, it would stand on its own two feet. It wouldn’t need circular financing and a cult of personality to prop it up, either. If it were truly special, there wouldn’t need to be an army of crazed acolytes that attack you for not pledging yourself to the graveyard smash. There has never been a tool or product in history sold with such hysteria and aggressive monocultural force that has ever turned out to be anything more than a grift. Some people have developed unhealthy relationships with large language models (LLMs) and the companies that make them, and that, not any certainty or proof of Artificial General Intelligence (AGI), is what motivates them.  This software is uniquely dark, both in what it unlocks in some people through its use and in the sense of the entities that sell it. Some people are in genuine awe of each of the rotation of clammy, soulless pod-people that saunter out of Anthropic every few weeks. Each one sounds a little weirder, more cultish, more disconnected from the real world. Silicon Valley may believe itself atheistic, but Anthropic has a worrying sense of fanaticism, both in the people that work there and its fanbase. Imagine the absolute worst fanbase of a video game possible, and then add layers of financialization, grifting and high school drama laced with pseudo-religious attachment. All for a fucking app!  Please, people. Nobody in the real world cares about “loops.” Nobody is thinking about tokenization. If you said inference to a guy on the street they’d take you to see a doctor. Nobody gives a shit. They don’t know what OpenClaw is either. Grow up. Go outside. You sound like a lunatic. Does your mother know how many Claude 20x accounts you have? It’s obsessive!  Anyway, the only reason that AI has any presence in our economy is that Microsoft, Google, Meta, and Amazon are intent on spending more than $765 billion in capital expenditures in 2026 and a trillion more in 2027 because they have no other hypergrowth ideas, even though generative AI has yet to show any real potential as something that can drive meaningful revenues (let alone profits), as evidenced by the fact that none of these companies break out their actual AI revenues , a point I made on CNBC late last week .  Google does not have the next Google Search, Microsoft does not have the next Microsoft Office, Meta does not have the next Facebook, and Amazon does not have the new AWS. That’s why they need you to believe that AI is a big deal without them ever having to prove why outside of capital expenditures. They want you to assume that all this money can’t be wrong , even though when you remove OpenAI and Anthropic ( who represent 89% of the revenues of the largest AI companies ) the AI industry is, at best, pulling in $20 billion in annual revenue. And lord do they want you to say “it’s early,” and that it’s just like the Dot Com Bubble , all so that you’ll either accept AI as your lord and savior or, alternatively, help justify one of the largest misallocations of capital in history as “building useful infrastructure.” Newsflash! AI GPUs are useful for generative AI and not much else. Every “innovation” in LLMs has only been made possible by throwing billions of dollars at the problem either in headcount or compute costs — every ounce of talent in the tech industry, every bit of media attention, every dollar of capital expenditures, all focused on one industry that has successfully created LLMs that are more expensive and significantly less useful than human beings .  The reason every AI person speaks in pie-in-the-sky hypotheticals is that the actual outcomes are decidedly mediocre when you compare them to their ruinous costs. Anthropic and OpenAI raised (assuming the rounds completely close) over $300 billion in 2026 alone, and take up the vast majority of available AI compute. They need you to speak in the future tense, because nothing — absolutely nothing — about what’s been created so far justifies even a fraction of its financial and infrastructural cost. When the AI bubble bursts, none of this infrastructure will be particularly useful. As I said in my premium about how this is worse than the Dot Com Bubble , GPUs are not fiber optic cable , and when the bubble bursts, NVIDIA chips will either be sitting in the coffers of the largest tech companies in the world, held by asset managers, or auctioned at a steep discount by creditors. These are not going to be useful for hobbyists, nor will they be cheaper to run, nor will incomplete data centers be cheaper to finish. The Dot Com era fiber overbuild was a result of a complete misread of demand signals, per Justin Kollar : It’s tempting to compare this to GPUs, but it doesn’t make sense at all!   You see, internet demand was a result of people wanting to get online and use the internet, with the leftover “useful infrastructure” having a blatantly obvious use case after the bubble burst, albeit one that took a lot longer to arrive than investors had hoped. There was no question about how that gear might be used or for what purpose one used fiber optic internet or networking gear, nor was there any question as to the underlying business model of offering an internet connection might mean.  We were also fairly early, and internet speeds were atrocious. In 2000 , only 52% of American adults were using the internet, and by 2003, that number had only increased to 61%. Per the World Bank , in 2005 only 16% of the world used the internet, and in 2024, that number had increased to 71%. When the internet was connected to via a 56k modem, access was charged by-the-minute, and obviously much, much slower than even the primitive (though expensive) broadband connections of the day.  While we’re used to connecting at speeds that make using a web-based app near-indistinguishable from one that runs on our computer, back in 2000, 2001, or 2002, the average US internet speed was, at best, 400 Kilobits/s , or roughly 50 kilobytes a second, compared to the average US internet speed of over 200 Megabits per second , or 25 megabytes a second.  Generative AI, on the other hand, is fucking everywhere , and anyone with an internet connection experiences it in effectively the same way. It’s non-consensually available in effectively every app — every Facebook, Google and Microsoft account, for example — and every media outlet known to man has mentioned AI multiple times since 2023. OpenAI and Anthropic might claim they need more data centers, but it’s unclear what “more data centers” actually achieves other than propping up NVIDIA and giving hyperscalers something to invest in.  A lack of data center capacity isn’t holding back people from using generative AI, nor is it stopping anybody from launching a product, nor can anyone actually express what it is that they’re being built for other than “reasons for Anthropic and OpenAI to spend money.” Anthropic’s supposed lack of compute did not stop it training or launching Mythos or Fable, and when it bought hundreds of megawatts of compute from SpaceX , the biggest news was that it expanded rate limits to allow users to burn $8,000 worth of tokens for $200 a month . Nothing about the painfully slow pace of data center development appears to be restraining a single AI company, outside of hyperscalers complaining they could’ve made more money from either Anthropic or Meta . In fact, the entire argument for more data centers appears to be “we need more compute so that people can buy it” far more than any cogent position around what these capacity shortages actually mean.  Who are the companies lining up to spend billions of dollars of compute — or, to be more specific, spend $435 billion or more to justify the $1 trillion in GPU sales that NVIDIA claims it’ll have by the end of 2027 ? That’s how much demand we’ll need. As NVIDIA intends to sell over a trillion dollars of Blackwell and Vera Rubin GPUs by the end of 2027 , it needs to have around (assuming a PUE of 1.35) 40GW of data center capacity built to support the 30GW+ of GPUs it will have sold . At about $12 a megawatt of critical IT (IE: the stuff in the data center that runs AI compute, and not everything else, like the cooling systems and any transmission loss), that’s $435 billion.  OpenAI estimates it’ll spend $50 billion on compute in 2026 , and Anthropic will likely spend comparable amounts. Otherwise, the only other player — outside of Microsoft, Google, and Amazon renting ( or backstopping ) capacity for Anthropic and OpenAI — with any meaningful compute spend is Meta (with Nebius and CoreWeave )... and Bloomberg is reporting that Meta is planning to start selling its compute because it doesn’t need all of it .  You’ll be shocked to hear that it might be renting some of that capacity… to Anthropic . Now NVIDIA is agreeing to financially backstop young cloud providers buying their GPUs by promising to rent back any unused capacity, yet another sign that actual, real demand does not exist at scale . AI boosters with black mold problems will say “this is just to help them raise debt,” to which I say “If the demand actually existed in any provable way, NVIDIA wouldn’t have to pay its customers to buy its products!”  Anyway, my larger point is that there was real demand during the dot com bubble, and LLMs’ demand appears decidedly artificial outside of OpenAI and Anthropic, who cannot afford to pay without unlimited venture capital funding.  This shit isn’t going to become magically cheaper once the bubble bursts, and considering the demand doesn’t appear to be there at scale with two-thirds of all venture capital funding focused on AI , I’m not sure what people expect to happen. Right now is the number one time in history where we should see near-infinite demand for compute across every single surface, and way more deals for compute capacity for companies other than the same four or five companies. Right now, as I’ve discussed before , Anthropic and OpenAI take up the majority of compute, leaving the rest of the world to fight for the leftover scraps, and because data centers take 18 to 36 months to build , capacity is taking forever to come online to fill the indeterminately-large amount of demand that remains. Nevertheless, said demand can’t be that large, otherwise we’d A) have other companies trying to build their own compute (other than Poolside, which failed to raise money to do so ) and B) massive remaining performance obligations — hundreds of billions of dollars’ worth — rather than the grim truth that 50% of hyperscaler RPOs are from Anthropic and OpenAI , inflating obligations by $448 billion, hiding the fact that Microsoft’s RPO growth is flat year-over-year and Amazon’s is only growing at a modest 20% when you remove Anthropic and OpenAI’s hundreds of billions of dollars’ of compute spend. Google’s is a little messier, as it’s hard to parse exactly how large its deals with Anthropic are thanks to its backstops and circular deals around Anthropic and its TPU chips . There’s also the compelling question as to what it is that anyone would be picking up once the bubble bursts. Demand for AI services is a direct result of the entire media, tech industry and venture capital ecosystem manufacturing consent for the use of LLMs, forcing them into every corner of every experience, something that will most decidedly end once the stock market and investors cease incentivizing it.  Once every media story isn’t about AI, once every Business Idiot with AI psychosis stops posting about it every day, when everyone stops asking about your AI strategy or wanking on about “sovereign AI,” it’ll become blatantly obvious that the actual demand for AI was not particularly strong. We have little compelling evidence that providing any inference-based services is profitable, which means that even if open source AI outlives the frontier AI labs, it’s unclear who would actually power the infrastructure. People can come up with however many weird blogs where they’ve done some napkin maths to try and extrapolate a potentially profitable inference provider, but I’ll only believe that one is profitable when someone shows me some fucking profit. And to be clear, without that profit, it’s unclear why anyone would offer these services at all. When you rent out a GPU cluster, you do so based on anticipated demand and the quality of service you want to provide. If you order too much, you’ve got a bunch of fallow capacity you’re paying for (and will lose money on), and if you order too little, you’ll have either unstable services or money left on the table…and even then, it’s unclear how profitable that would be.  AI demand is, at this point, a direct result of societal pressure and non-consensually overwhelming customers with AI features. While there are people that like and pay for ChatGPT or Claude, those who do so on a subscription basis are doing so because they can get $30 to $40 of compute for a dollar . The vast, vast majority of AI compute demand is from services provided to people either for free or sold at such a massive discount that it’s impossible that anyone on a $20 or $200-a-month plan could even afford these services had they paid their actual token cost. To paraphrase Cory Doctorow, your demand is based on selling $40 for a dollar. That’s not a real business, nor is that organic demand. One could argue that “these services will become cheaper,” but that would require them to… become cheaper. More compute isn’t (and hasn’t) lowering the cost of AI. Newer GPUs aren’t lowering the cost. Barely-tested Broadcom GPUs , Amazon Trainium XPUs, and Google TPUs aren’t lowering the costs. Even if they were to somehow magically do so in the future, what do we do with the H100, H200, B100, B200, B300 or AMD GPUs? Melt them down for scrap? Steal the RAM? Build a GPU fort?  The Dot Com (and, by extension, telecom) Bubble was never a question of whether the internet was a useful thing that people would pay for , nor were there journalists and dodgy studies that desperately pleaded with us that AI is here, and it’s real.  Everybody has access to AI now! They can all see it and use it if they want to, and they’ve got lots and lots of ways to pay for it! Maybe the reason that AI revenues are so putrid is that they don’t really have any reasons to pay for it, either because the free services do most of what they need (IE: google searches) or subsidized subscriptions that cost $200 a month allow them to burn as much compute whipping up HTML-based calorie tracking apps that get two users. Every time I read somebody on Twitter say that “we’re early” or that “most people haven’t even tried agents” I feel like screaming. Motherfucker, everyone is talking about agents in every single media property all the time . AI boosters will refer to literally any AI feature as an agent, even if it’s a basic web search or generating code. The reason that most people are kind of “meh” about AI is that it doesn’t do things that they associate with AI (autonomously and automatically taking care of the things they need with little prompting or coaxing), everybody knows it hallucinates, and AI data centers are horrifying monoliths of capital that get massive tax breaks, use a ton of water , belch toxins into the air , and are being built by faceless corporations, ultra-oafs like Kevin “Mr. Dogshit” O’leary , or charmlessly damp Valley elitists like Altman and Amodei. Every single person freaking out about “what if China does AI better than America” is living in a child’s fantasy. Oh no! China might get Mythos-level AI? Bad news folks! Anthropic itself already admitted that cheaper models — including Claude Haiku 4.5 and Kimi K2.7 — were able to identify the very same vulnerabilities as Fable (so, Mythos with guardrails).  China has cheap power, data center capacity, and NVIDIA’s Blackwell GPUs . The thing that everybody is scared of has happened already, and you know what else happened? Nothing, because they, like American AI labs, are building LLMs. The only thing that American labs are scared of is cheaper open source Chinese models offering similar performance to their premium products , something that has also already happened.  Remember: the only people that can afford to build data centers are either hyperscalers ( that are now having to fund the buildout with debt as their cash flow turns negative ), Oracle ( which will die if OpenAI can’t pay it ), unprofitable neoclouds , and land speculators. AI data centers are massive, expensive operations, and raising money to finish (or furnish) one after the bubble bursts will be very, very difficult. I realize that everybody wants there to be a happy ending after all of this collapses. I get that it’s easier to think of things in familiar terms — even if said terms involved a 77% drop in the NASDAQ — because there was something good and nice at the end. But doing so only serves to help protect the interests — and brands! — of venture capitalists, asset managers, private credit funds , hyperscalers, captured tech and business journalists and sell-side analysts that insisted on ignoring every warning sign and waving away problems by saying it was “just like Uber ( nope !)” or “just like Amazon Web Services ( between 2003 and 2015, Amazon spent $29.7 billion on capex, normalized for inflation ),” or simply saying that “yes it’s a bubble, but bubbles lead to great industries.” GPUs aren’t dark fiber! GPUs aren’t fucking railroads! GPUs are GPUs! They are used for basically one thing ! And that one thing lacks meaningful demand outside of subsidized services and circular financing!  And now people are discussing a bailout like this is 2008, and I must be clear how different this is, and how little it resembles the Great Financial Crisis! The AI industry has demanded everything from us — more money than has ever been invested, more power than anything has ever needed, the stolen works of millions of hard-working creatives , so many GPUs and so many data centers that it’s causing a global supply chain crisis and a new class of RAM and storage-based inflation , the majority of venture capital funding ,  and constant attention focused on an endless campaign of fear-mongering with the express intention of hyping a technology based on a mixture of mysticism and outright lies — and still, even as we enter the late innings of the bubble, it wants more.  Capital-hog Sam Altman has floated the idea of handing 5% of OpenAI to the US government , a stake worth around $42 billion, claiming that (to quote the FT) “...giving the public a financial stake in the company is the best way to share the upside of AI,” failing to note what said upside might be, likely because there isn’t one unless “the public” refers to “the shareholders of OpenAI.”  It isn’t clear how this would happen, outside of it requiring congressional approval as a result of the Takings Clause of the Fifth Amendment , which states that “private property [can’t] be taken for public use without just compensation,” meaning that the US government would likely have to buy the stock at whatever valuation it considered “just.”  Yet the FT had one other interesting tidbit — that Altman is suggesting that whatever this is would “...would involve other US AI companies handing over a similar stake, although it is not clear if the other labs would be willing to do so”: This is, just to be clear, not a bailout. Even though it’s blatantly obvious that Altman wants to cozy up to the Trump Administration and, he hopes, get $42 billion of funding to attach his questionably-valued quasi-startup, $42 billion is $8 billion less than OpenAI will spend on compute in 2026 , and considering OpenAI has projected to burn $852 billion through the end of 2030 , that 5% stake would only exist to prolong the inevitable. You see, a bailout usually has an endpoint — a time at which the company in question no longer needs the funds.  So, let’s be clear about something : we’re actually in several bubbles at once. The great financial crisis, by comparison, was two major bubbles (per my piece on how AI Isn’t Too Big To Fail from a few months ago) — the over-investment and speculation on mortgages (both subprime and otherwise), and the collapse of the commercial paper (a type of loan) market that kept much of the banking system functioning, which was the real “Too Big To Fail”: Commercial paper was, at the time, often paid off using more commercial paper, and when AIG’s credit rating dropped in the middle of September 2008 , it was unable to roll over its debt (by which I mean “get new commercial paper to pay off its old commercial paper”), and money market funds like Fidelity couldn’t even buy it anymore because it wasn’t investment grade, which meant that AIG couldn’t pay back its loans.  While I won’t recount the entirety of the premium (mostly because it’s super long), AIG was deemed “Too Big To Fail” because it would’ve exploded the markets had it done so. Michael Lewitt, an economist and money manager, described a hypothetical AIG failure as being “as close to an extinction-level event as the financial markets have seen since the Great Depression” in a New York Times op-ed: Yet the real “Too Big To Fail” was far quieter and more malignant, taking the form of trillions of dollars funnelled to banks: The banking system ran (and still runs) on overnight facilities like the federal repo market, where financial institutions offer up collateral — like, say, mortgages — as a means of funding their day-to-day operations. Previously, money market funds were the lenders in the repo market…except they were now a little hesitant to take that collateral, which forced the government to step in with the PDCF (which traded risky, frozen assets like subprime mortgages for cash to avoid a default) and the TSLF (which traded risky bonds for US treasuries). Absolutely nothing about these facilities or anything to do with “too big to fail” were to do with stabilizing the stock market, which was effectively cut in half , with unemployment spiking to 10% . These measures existed exclusively to protect the financial system, with only $46 billion (about 10%) focused on trying to save homeowners from foreclosure , and in the end, to quote a congressional panel from 2009 , “...the panel sees no evidence that Treasury has used TARP funds to support the housing market by avoiding preventable foreclosures.”  The Troubled Asset Relief Program (TARP) spent over $400 billion to bail out the banks, financial institutions and auto industry that would’ve collapsed as a result of an economy-wide lending freeze. Nobody went to jail, nothing really changed, and banks still don’t have to keep reserves thanks to changes made around COVID. By comparison, OpenAI and Anthropic are systemically irrelevant, much like the rest of the generative AI industry. While their existence supports the overall symbolic value of the US stock market, their actual economic presence is minor, outside of what I estimate is around $75 billion to $100 billion of 2026 compute spend and what will likely be around $60 billion of combined revenue, with the rest of the AI industry having so little that it’s barely worth thinking about. It’s also unclear what you’d bail out, unless the plan is to feed them capital for all eternity until they work out how to run a functional business (so, forever). Neither of them have significant debt — and Broadcom is backstopping $30 billion of Anthropic’s $35 billion TPU deal with Apollo — and their equity positions (outside of SoftBank, which I’ll get to) are only load-bearing to venture capitalists in the sense that their fund vintages will painfully sour if they’re unable to go public.  There is no avoiding the carnage to come, outside of there being somewhere in the order of ten to a hundred times the demand for AI compute by 2030 that exists today, which would require AI compute to be larger than the $779 billion that the software industry earns annually .  There is no bailout that can reverse the trend once demand wanes for NVIDIA’s GPUs after hyperscalers reduce their capex, which will in turn kill the revenues of Taiwanese ODMs that build AI servers for hyperscalers , which will in turn kill the revenues of RAM and storage companies, which will lead to a prolonged depression throughout a semiconductor industry addicted to hopium peddled by a tech industry ruled by Business Idiots that have no idea what to do other than hire people, fire people and spend money .  As I’ve said many times, people are conflating massive capital expenditures — invested through debt-fueled data center speculation and hyperscalers bereft of hypergrowth ideas — with real, diverse and consistent AI demand, pumping valuations based on vibes rather than reality , which means that when vibes take a violent, permanent shift, nobody has anything to point to as a means of turning people’s frowns upside down. The collapse in value of AI startups wouldn’t be changed by a bailout unless the US government literally invested in worthless startups as a means of propping up venture capital, and said “bailout” would number in the hundreds of billions of dollars, and while I know you’re gonna say “ohhhh Trump is so corrupt oooh Trump will do this Trump will do that,” this is not a rational or logical or even historically-accurate thing to say.  Trump cannot simply mobilize $50 billion or $100 billion. It will go through the House and the Senate, and any bailout of the AI sector would be an incredibly-unpopular decision, infuriating not just those on the left who’ve grown tired of Big Tech, but with those Republicans that pretend to care about working Americans or fiscal probity.  As a reminder, the first vote of the 2008 bailout failed, with Republicans and Democrats each fairly split on how they felt about the bill — and that rejection happened during a time when the US financial system was quite literally falling to shit.  As far as the data center bubble goes, the government is absolutely willing to let unfinished or abandoned properties lay dormant. In the final quarter of 2008, 11% of US homes were empty , or 15% if you include vacation homes.  Banks that have invested in data centers that have yet to be built (or start construction) can (and will) resell the land, though likely at a loss, and land retains value even if you haven’t built a giant warehouse full of GPUs that only lose money. There isn’t a need for a bailout here, and one won’t be forthcoming. After the Global Financial Crisis, builders were allowed to collapse to the extent that the number of construction firms halved in America between 2007 and 2012 . You could argue that Trump “will just do that this time,” or that he’ll “get a bribe” or something, but is that really the best you’ve got? Scary stories about the President? If every answer you have is “but Trump will just do it,” you’re not analyzing, you’re catastrophizing.  And, most crucially, the vast majority of big tech will be fine, at least in the short term, when the bubble bursts. NVIDIA will likely cease being the largest company on the stock market, and the Magnificent Seven will have a dramatic fall from grace, but outside of unforeseen horrendous financial decisions, the worst I could see would be impairments for Microsoft, Google, Meta, and Amazon, and SEC action against NVIDIA if it did actually sell GPUs to China. This doesn’t mean that things won’t fucking suck for anyone in the market, nor that the vast majority of people won’t fucking suffer as they always do when bubbles burst.  Which is why I am making a firm, clear statement to end this piece. I repeat myself: No bailouts, no handouts, no special treatment, no tax breaks, no CHIPS act, and no sovereign wealth fund. It is time to tell the AI industry to go fuck itself, because it’s effectively done the same to the rest of society. These companies must be forced to stand on their own two feet and die with dignity if their wretched business models can’t keep up. The world’s governments have rolled on their backs and shown their bellies to the tech industry for far too long, and have been aggressively conned by some of the richest people alive into believing that fucking Sam Altman and Dario Amodei are building anything other than the world’s least-profitable software.  We do not need a “sovereign AI strategy,” nor do we need “a sovereign AI wealth fund,” nor do we need to “make sure America leads in AI,” at least not when we’re talking about large language models, the underlying technology of ChatGPT and Claude, two of the most over-hyped and deceptively-marketed pieces of software in history.  Whether or not LLMs are a useful tool is irrelevant, because the AI industry has demanded the world hand it as much land and money and as many resources as it desires to continue proliferating a technology that has only ever lost money and has no path to sustainability. The only reason it has gone anywhere is because the tech industry has united around it as a means of hiding from the fact it has no next big thing , and nothing — absolutely nothing — that a LLM can do remotely justifies the investment. And it has only got this far because of a captured business and tech media overstating its capabilities and hand-waving its obvious efficacy issues and economic instability. There are too many that have proven easily-wooed by whimsical white boys that promise they’re building machine intelligence, and when the markets bleed red, these people should know that they’re responsible. So much of the so-called journalism around AI has been used to enrich the already-rich and inflate a bubble that will hurt hundreds of millions of regular people globally as Sam Altman and Dario Amodei remain billionaires despite their companies’ fates. When the time comes, the AI industry must burn. It must be allowed to die. Generative AI has already been given far too much money, oxygen and attention, and if it cannot survive without continual venture capital and media coddling, it is unworthy and unnecessary, and must face the cold, hard reality that every regular person faces when they fail. And there is no “bailing out” these wretched firms. Giving $42 billion to OpenAI or Anthropic will not fix their business models, nor will it magic up the $400 billion or more in annual revenue to substantiate just NVIDIA’s AI GPU sales through the end of 2027.   These people are not building the future — they’re finding ways to re-entrench the status quo, to give Microsoft, Google, Amazon and Meta ways to grow their revenues and centralize infrastructure under the auspices of “innovation.”  If any policy makers read this, know that you’ve been had by the AI industry. They want you to believe they’re essential so you’ll bail them and their rich friends out when the time comes, or funnel taxpayer funds into building them data centers. They are not building autonomous intelligence, nor will they ever do so.  I think it’s fanciful to imagine that there would ever be actual consequences for this bubble, but if there are, the people to hold responsible are Sam Altman, Dario Amodei, Satya Nadella, Sundar Pichai, Andy Jassy, Jensen Huang, Mark Zuckerberg, and everyone else who forcefully manufactured consent for a dead end technology and built the rails to serve the world its next great financial crisis. Until something changes, the tech industry will never be capable of building anything other than consensus and reinforcements of the status quo. So, spit in the face of those who even hint at a bailout, refuse to accept it, and demand that they do the complex, ugly work of thinking about the actual consequences of everyone being wrong. When this era ends, we will need to thoroughly excavate the collapse to make sure it doesn’t happen again, identifying the organizations and personalities that were used to manufacture consent and spread mythology about LLMs.  Every major bubble that has ever happened has mostly left the stones of responsibility unturned. The carnage that I fear will follow this era’s collapse will be horrifying, and we must do everything in our power to both thoroughly understand how we got here and make sure it doesn’t happen again, which will involve many hard conversations about our financial system, media ecosystem, and how innovation is invested in, built, bought and sold.  The same goes for the acolytes of this era. There are people who have developed a genuine hostility toward those who do not immediately accept a for-profit entity as their lord and savior. This is a sickness within the tech industry that must be put to an end.  Much of this will be unavoidable, because I think what follows the AI bubble will be a greater revaluation of the tech industry, a necessary reckoning with reality for a Silicon Valley that’s far more beholden to capital than it is human progress. The cults of personality that dominate this industry do not care about you, or me, or anyone other than those they revere and their theoretical placement in their dream of a society dominated by the rich and their chosen cronies. I refuse to accept their future as an inevitability. As I said a few weeks ago: This era must end, and all failures must be allowed to fail.  Let AI burn. If you liked this piece, you should subscribe to my premium newsletter. It’s $70 a year, or $7 a month, and in return you get a weekly newsletter that’s usually anywhere from 10,000 to 18,000 words, including vast, detailed analyses of the biggest events and companies in the AI bubble.  The stock market bubble, where both the value of stocks and the earnings of companies in the market are inflated to an historic level . A data center speculation bubble, where I believe we’re building AI GPU capacity in expectation of $450 billion or more in annual data center revenue for an industry that, without two unsustainable venture-backed oafs, has a few billion dollars’ worth of demand. An AI startup bubble, where the vast majority of AI startups are both over-valued and have no foreseeable path to acquisition or a public offering . These startups also rely on buying tokens from OpenAI and Anthropic, making them far more cash-intensive, making them absorb the majority of venture capital funding. A private credit bubble, where asset managers have sunk billions of dollars of pension and insurance funds into AI data centers .  A semiconductor bubble, where supply chains have become saturated with demand from those building AI data centers, inflating the cost of RAM and storage , making all electronics more expensive, including those inside the AI data centers, creating a vicious cycle that has doubled the cost of a gigawatt data center from $50 billion to $100 billion in a little under 10 months.

0 views
Jim Nielsen 1 months ago

Making a Shuffle Button

I made some updates to my notes blog , including a change to how my “Shuffle” feature worked. Figured I’d blog about it. At the time of this writing, I have 974 “notes” that I’ve published. For fun, I have a “shuffle” button that digs up a random note from the past. I like to press it from time to time and re-encounter some insight from the past. It’s like going through an old album, pulling out a random photo, and thinking, “Oh yeah, I remember this! Good times.” Like old photos, there’s also the occasional “that didn’t age so well”. But I find it fun to randomly dig up old insights from others and continue to be inspired. Since my site is built and hosted as static files without a runtime server, this feature required JavaScript to work. Every page had a snippet like this: Essentially: inject every note ID into every HTML page and, when the shuffle button is clicked, randomly grab one and navigate the user to it. Not the most elegant thing, but it worked. The problem was that every time I published a new post, every single page had to be re-uploaded to Netlify because every file’s hash would change and its etag/cache was invalidated. This made my builds slow. It also made it difficult, from a development perspective, to ensure refactors didn’t result in unexpected changes to output (using from my SSG web origami ). So I decided to make a change. Because I love to see if I can make things work without JavaScript, I had the thought to randomly write the at build time using my SSG, which would result in output like this: And every time I re-build my site, just have this logic run on the static site generator so that it’s different for every page, every time. I decided I didn’t want to do this, so on to JavaScript! My first thought was to create a single JSON file that contained all my note IDs. Then when the “Shuffle” button gets clicked, I fetch that, grab a random ID, and navigate the user, e.g. This would work. It localizes the caching issue to a single file, so only one file has to be invalidated/re-uploaded across builds. But in playing with it a little more, I decided to try something a little more...unconventional. I’ve written before about having lots of little HTML pages and I thought, “Can I put this functionality in a single HTML page rather than a JSON file?” And what I ended up with was a link, e.g. That when clicked navigates the user to a new page. That page has all the JS logic embedded in it, e.g. There are a few things I like about the experience this implementation provides. First: shuffle is a route , so I can navigate to it directly without using the GUI, e.g. notes.jim-nielsen.com/shuffle Second: I handle the UI/X with a slight delay to make it appear like something is happening when you click the button. If you click the button and it immediately jumps to the next, randomized page, it almost seems to happen too fast. Like you’re left with this feeling of “What just happened?” But in this scenario, it navigates you to the “Shuffle” page, the button you just clicked turns into a spinner + text indicating something is happening, and there’s a slight (intentional) delay before the JS executes and sends you to a randomized note. I know it’s a bit weird. “Introduce artificial slowness? Are you crazy?” But I like it. It feels like the shuffle feature on an old music player. I remember one of my CD players had a “Shuffle” feature. When I’d click the button, it would display “Shuffling…” on the little black and white screen and you’d encounter this brief state where (I presume) the lens inside the hardware would move along the physical track to the spot where it would start reading a new, random song from the CD. The hardware constraints necessitated this kind of an experience, but I always liked it because it felt like the CD player was “thinking” about what track to pick next. This state clearly conveyed to me that my intent to shuffle was received and being followed. I liked that feedback, and it’s exactly what I wanted to do on my notes site (even though it was completely unnecessary). I like having that brief moment of feedback where it’s very clear that your intention was received and being followed, vs. having it happen so fast you can’t even perceive precisely what happened. Here’s a video to show it in action: I know that’s a lot of information for something so small — and, arguably, unnecessary. But I still enjoy writing about how I make decisions when I build things for myself. Hence this post. Reply via: Email · Mastodon · Bluesky Doesn’t require JavaScript Doesn’t require a server (request-time logic) File hashes change across builds (even if there’s no new content or template changes, every HTML page now has a different for the shuffle link for every build ). This makes deployments way slower because Netlify has to redeploy every file on every build. Plus Etags change so caching is basically ineffectual.

0 views
The Jolly Teapot 1 months ago

A peculiar bug in Safari

On weekend mornings, I have the inescapable habit of looking at my website and seeing what I can change, what I can remove, what I can improve in terms of HTML, CSS, layout, links, etc. This Saturday, as I wanted to look closer at the way the period at the end of a sentence rendered when appearing just after a word in italic (I know), I noticed something curious. When I zoomed in the page, using “Command – Plus Sign” (⌘+), I could see that the line length was changing with the size of the text. The bigger the text, the longer the line. You see, I’m very protective of the I use on this site —  — especially for Mac users, who see it in the Charter font. *1 This value sets an ideal number of characters for each line making it, when paired with the right line height, easier to read (supposedly). Zooming in on text shouldn’t change the line length, so I looked around and realised that I was a bit clueless when it comes to identifying bugs, and even checking if they were already reported. I found a few bug reports related to zooming in, but none of them described my issue. Not only that, but I didn’t really know if this was a Webkit problem, or a Safari problem. So instead of working my way to either confirming an existing bug or filing a new one , I did what I usually do when facing a problem: I avoided it altogether rather than trying to solve it. Therefore I changed to in my CSS, resulting in a similar line length for Charter. *2 With as the unit, zooming doesn’t modify the line length, so I’m pretty happy with this easy fix. Bonus point: takes up the same number of bytes as in my default CSS, still capped at 132 bytes. Imagine the extra-byte horror if I had to use something like or ? It would have ruined my sunny Saturday morning. This little website update made me realise something: my site design is pretty much done, and I hadn’t changed anything for a few weeks or even months. I actually miss the satisfaction of changing something at the end of my little routine. Checking every detail on every page, revisiting every line of code just to see what can be improved, even if it’s just removing extra quotation marks in an attribute or an optional closing tag, is not as fun when there is nothing to do at the end. I really like my site’s current design, and even if there might be a few tiny tweaks like this one in the future, I feel that the overall look and feel is pretty much final. It’s a weird feeling, but now I have no excuse for not writing more, and publishing more posts, even if they are unfinished , or shorter than usual . For others, falling back to the default serif, usually Times New Roman, is indeed a bit narrow; or would be better, but it’s too wide for Charter.  ^ For the serif/Times New Roman fallback, creates a slightly longer line, which is atually better than what it was with .  ^ For others, falling back to the default serif, usually Times New Roman, is indeed a bit narrow; or would be better, but it’s too wide for Charter.  ^ For the serif/Times New Roman fallback, creates a slightly longer line, which is atually better than what it was with .  ^

0 views