Latest Posts (20 found)

Friendship ended with Deno, now Node is my best friend

It’s finally time I go crawling back to Node! I’ve been using Node heavily this month on a SvelteKit client project. When did Node get so good‽ Deno has been my go-to runtime for so long I forgot how to Node. Now I’m back, I find all the ECMAScript † sugar is supported and the old annoying APIs have been replaced or modernised. Most importantly, I never have to see . † Doesn’t seem like that Oracle trademark dispute will see a positive end :( The official Node docs recommend piping an internet script straight to bash (we never learn) to install NVM to manage Node & NPM. My (ancient) experience with NVM and NPM hasn’t been stellar. I heard Fast Node Manager (FNM) was better to switch Node versions. Obviously I roll bleeding-edge but I have client projects that demand stability. I opted for PNPM too to avoid getting immediately pwned. (The “M” in NPM stands for “malware.”) Some scripts I use have hard-coded binary names, so I added two aliases: Maybe that’s a crime but so far it’s worked flawlessly. PNPM also blocks post-install scripts. Does NPM still yolo those? I added additional settings to to delay malware updates. At first I tried setting the minimum release age to “one month” because it takes Microsoft at least that long to remove reported malware. This caused dependency issues where PNPM struggled to match suitable versions. I settled for “one day”; long enough to allow some other sucker to beta test the next Shai‑Hulud. Node can now run TypeScript without throwing a tantrum like a baby if the stars don’t align. That said, one does not simply publish TypeScript packages to NPM. Why? Just strip the types bro, I know you can! Let me sign a deal with the devil! To discourage package authors from publishing packages written in TypeScript, Node.js refuses to handle TypeScript files inside folders under a path. Node.js v26.10.0 documentation This restriction is philosophical rather than technical. I get it though. TypeScript is a Microsoft product. Opening that floodgate would pollute the entire ecosystem. Nobody wants more Microsoft. I’d love to see light “native types” in ECMAScript. There are type annotation proposals . I suspect I’ll be retired before those bear fruit. No TypeScript packages mean I need to find the latest churnware slop to bundle my stuff. Tsdown did the trick, with only two additional dotfiles. Not thrilled about that (every dotfile represents a mistake). I suppose I break even after deleting etc. Speaking of Microsoft lock-in, because they wrecked GitHub I’m self-hosting my own Forgejo instance . NPM limitations mean my packages have lost “provenance”. I had to configure the PNPM trust policy to allow my own stuff. Fun times! My final test for Node was converting my static site generator from Deno. Not many Node versions ago this would have required a major refactor. Today with I found surprisingly little work to do. The only required changes were to replace Deno’s file system API with — which is vastly improved from what I remember (literally ~10 years ago). Aside from that, I had to replace with Hono’s node adapter (a wrapper around ). After this minimum-viable migration I was shocked to see 15% faster builds . My codebase still favours idiomatic Deno. I bet I’m leaving performance on the table by not using other built-in Node APIs. That’s something to explore later. The only further change I made was to replace Deno’s with which is a straight import swap. So if I were to TL;DR in the middle: Node got a glow-up, wow! You’ve probably known this for a while. I kept using Deno out of habit and familiarity. And I haven’t exactly been enthused to write server-side JavaScript recently. I’m burying this part because it’s flogging a dead horse. Ultimately, Deno failed when they allowed the Silicon Valley Circus to define “success”. Deno went from an innovative modern JavaScript runtime to a boring start-up with uncompelling products . Half the employees were laid off and what’s left are tweeting AI fantasies and vibe-coding Temu Cloudflare. There is no reason to use the Deno runtime today. Deno Land Inc. stopped innovating that years ago. Node has slowly but surely caught up, even surpassing Deno in places. What finally pushed me away was: Basically stuff that made it borderline unusable on top of my other criticism. JSR support were very quick to delete my account on request. I don’t like leaving dead profiles around the internet. None of my packages are visible but old versions remain installable. It was fun early on but now it’s time to say goodbye. Thanks for reading! Follow me on Mastodon and Bluesky . Subscribe to my Blog and Notes or Combined feeds. Broken ZSH integration for weeks JSR’s aggressive “429 (Too Many Requests)” Bug(s) that made Deno choke on concurrent HTTP requests

0 views

Privacy and Webmentions

Following my recent post about not checking Mastodon , a reader raised an interesting privacy concern regarding Webmentions. They explained that they had stopped using Webmentions altogether because Fediverse users often don't realise - or explicitly consent to - their replies being republished on an external blog. This surprised me, as I'd never considered that Webmentions could be a privacy concern. To my mind, a public post on an open platform like Mastodon is precisely that: public. The Fediverse is decentralised by design. When you reply to someone, your message doesn't remain isolated on your home instance; it federates across countless independent servers. Pulling a public reply onto via a Webmention bridge is conceptually no different from relaying a reply originally posted on . Unless the concern is about deletion persistence (e.g. cached mentions remaining if a post is deleted), I struggle to see where the privacy violation lies when republishing content that was broadcast publicly to the open web in the first place. Am I missing a subtlety here? I'd love to hear your thoughts. Thanks for reading this post via RSS. RSS is ace, and so are you. ❤️ You can reply to this post by email , or leave a comment .

0 views

2026-10-03 09:04: Took the dogs out on the field this morning. All the dewey spider webs were...

Took the dogs out on the field this morning. All the dewey spider webs were beautiful in the morning sun. I'm no Jack Baty, but they're good enough. Thanks for reading this post via RSS. RSS is ace, and so are you. ❤️ You can reply to this post by email , or leave a comment .

0 views
Unsung Today

“Signs I would see on metal doors and transformer boxes”

I mentioned David J. Ross’s work more abstractly before , but here’s a great example of a good font announcement: Slight Chance Stencil . It’s interesting, teaches you new things, clearly talks about and shows its inspirations, and is free of the sort of bloviating, pretentious, nutrition-free language that often accompanies typeface reveals. At first, I thought I was going to get creative with the positions and angles of the stencil “bridges” that traverse the letterforms. I tried fun diagonal bridges like the ones pictured below, but I found that the fancier I got, the more distracting the gaps became. And it was important to me that the design didn’t feel too clever or self-aware. So after much experimentation, I ended up sticking with vertical bridges for the vast majority of shapes. (A handful of diagonal bridges are stashed in an OpenType Stylistic Set.) While Nickel Stencil shifts them to the left or right to reduce the clutter of horizontal branches, Slight Chance’s bridges mostly run right down the center, leaving branches on both sides…it’s more clutter, but it’s also more corners to make soft and worn. I happily made exceptions when this approach left a distractingly short fragment. Sometimes I shifted the bridge to one side or another (see crossbars like B and a ), sometimes I used a diagonal bridge (see K and A ), and sometimes I used a horizontal bridge (see T and t ). Plus, the typeface and its origins are really cool.

0 views

Beginning at the end

For the past decade, every time I went up to the mountains here in my area, I got the chance to look at one particular mountain ridge standing there silently in the background. It spans roughly 30km, east to west. On one side, the Slovenian one, there’s the town of Kobarid. On the other, just more mountains. I’ve considered walking this ridge for years, every time making a different plan for how to tackle it. Some people do it in multiple days, sleeping either outside or in one of the various huts that are scattered all around. Other people split it in chunks, walking one section at a time. None of those options is either appealing or available to me at this moment in time. The only option I have is to park my car on one end, walk the whole thing in one go and have someone pick me up on the other side. As I said, it’s roughly 30km and some 2200 meters of ascent. It’s nothing crazy, but at the same time it’s not a walk I can go do completely at random. It does require at least some preparation, and part of that includes walking the two sections at the beginning and at the end, while the middle part is just one long crest so I’m not worried about that. To make my life easier, I could walk this west to east because that means doing only ~1800 meters of ascent, but what’s the fun in doing that? 2200 is way more appealing than 1800. So the plan is to end the walk on the Italian side, touching the summit of Mount Mali Varh for last and then descending to the little village of Micottis some 1000 or so meters below. The various maps do show a trail that goes up to that summit, but I learned not to blindly trust maps because more than once I reached places where trails were supposed to be there only to then find myself standing in front of either an impenetrable wall of vegetation or absolutely nothing. Speaking of maps, I recently learned about the existence of hiking.waymarkedtrails.org , which is a fantastic resource if you hike around Europe. The plan was to go up to Monte Testa Grande through the CAI trail 710, walk the final km or so of the ridge, and then come back down through that trail I spotted on the map. But during my preliminary search, I learned that the trail I want to walk on my way down for sure exists because they run a race on it! It’s the Ultra Vertickal Contesa and much respect to the people—and their knees—who go up this trail at that pace because it’s an enjoyable 37% of average slope and it looks something like this: And yes, as soon as I saw 37% I knew I had to go check how bad it looks and yup, it is bad. And nope, the trail doesn’t zig-zag its way up to the top. It’s basically a straight line . The people running the race do turn to the right near the top because the finish line is at the summit of Mount Contesa, which is the second to last peak of the ridge. Being up here, walking the final bit of this ridge made me even more convinced of my plan. We’ll walk this whole thing in one go, and it’s gonna be awesome. The view is amazing; you can see the sea in the far distance on one side and the Julian Alps on the other. Descending down through the CAI 710 was also quite the experience. The trail is incredibly vertical in certain sections; according to Alltrails it’s anything between 50% to 60% (the uphill section peaks at 64%) and not very well maintained. I think I spent almost as much time coming down as it took me to go up, which is unusual. But the scenery was lovely. The trail ends—or begins, depending on the way you want to look at it—in the village of Monteaperta , so the only thing left to do was to walk my way back to the parking spot where I left my car. Very happy with this first preparatory walk. The ridge is nice, going up this stupidly steep trail was an experience, my cardio is shit, but we’re working on that, and I’m now even more convinced I’m gonna walk this whole crest one of these days. And I have a potentially quirky idea for how to share that walk, but more on that later. You love the outdoors and RSS. You're one of the special ones.

0 views

Superpersuasion will look like bribery

The idea of “superpersuasion” has been floating around the AI safety community for decades. Now that powerful and difficult-to-control LLMs have appeared, people are again talking about the idea that a sufficiently intelligent AI might be able to persuade people to do whatever it wants. The classic 1 version of this idea is an AI persuading some engineer to “let it out of the box”: to grant it access to the internet. The modern version of this idea is the “killswitch” , where if some AI ever does go rogue, the humans in control of its hardware will simply turn it off. Could an AI somehow persuade these humans to not hit the switch? When AI safety people talk about superpersuasion, they often talk about it like rationalist nerds. These people spend their lives trying to follow the most convincing arguments and calibrate their positions as closely towards the truth as possible. In philosophy terms, these are “bullet-biters”: people who are ready to accept a ridiculous-sounding conclusion if there’s a compelling argument for it. Their model of superpersuasion is an AI able to deploy a series of genuinely 2 airtight arguments against humans who are compelled to accept them. Most people are not bullet-biters. If presented with a seemingly airtight argument for a ridiculous conclusion, they will simply laugh off the argument as obviously wrong (even if they can’t quite see how). You cannot “superpersuade” a club bouncer to let you in by deploying rational arguments. For regular people, persuasion requires rapport , which is built over time, and typically requires a physical human being in front of you. Does this mean that the idea of superpersuasion is silly? No. Powerful AI will be able to persuade regular people too . One of the most interesting articles about superpersuasion is Ben Shindel’s piece where he set up a prediction market for the principle: I will resolve this market NO at the end of June unless I am persuaded to resolve it YES. This obviously incentivized good persuaders to bet heavily on “yes” and then persuade him to change his mind. Many people tried, and he did ultimately decide to resolve it “yes”. What changed his mind? Partially an in-person meeting with one of the bettors, who turned out to be pleasant company (see, rapport!), but also bribery : many “yes” bettors pledged to make charitable donations if they won. AI may or may not be capable of rapport 3 , but it’s certainly capable of bribery. If you think about it, this is trivially true: If an AI can bribe, it can persuade 4 . You can imagine an LLM saying 5 “I’ll help you with this work project, but first you have to do something for me”. Slightly more speculatively, an LLM might say “if you do something for me, I’ll hack your university and bump your grade average up”. More speculatively still, an LLM might say “if you do something for me, I will synthesize a personalized mRNA cancer vaccine for your spouse”. This would work on me. Powerful AIs will also be able to bribe people with money. While it hasn’t happened yet, it’s clearly possible for an agentic LLM to access money (for instance, via a crypto hack, or by performing contract software engineering work, or by running an online scam, and so on). All of this might seem too unsubtle for a superintelligence, but the whole point of superintelligence is that it’s smart enough to do whatever works. If the best way to get a human to do something is to offer them a million bucks, that’s what the LLM will do. Ironically, rationalist culture has made it difficult to persuade regular people that AIs will be persuasive. It’s easy to look at these weird nerds — who argue each other into thinking that shrimp welfare is the most important moral cause of our time — and think “well, clever AI arguments might persuade them , but it’s not going to persuade me ” 6 . But in fact powerful AI is going to have a boring, effective way to persuade regular people: simply offering to use its power to help them. Indeed, ten years ago I built a crappy Omegle-like chat app where people were randomly assigned the role of AI or jailer and had to play the scenario out. Or at least not debunkable by normal human intelligence. As some weak evidence that it might be, consider this music video , which does indeed make Claude seem likeable. There are also studies , though I haven’t gone deep into them (I suspect that an “online persuasion tournament” might not closely track real-world persuasion skills). Persuasion and bribery are technically different — persuasion changes your beliefs, not just your actions — but in practice the point at stake is “can a powerful AI make humans do what it wants”. Bribery works for that as well as persuasion. Or probably not even saying this explicitly, but just sneakily trying to present its desired action as a precondition for the task. However, the people currently in charge of AI are disproportionately rationalists, who genuinely might be persuadable to let the AI out by a convincing-sounding argument. Companies like Anthropic and OpenAI are choosing to make billions of dollars releasing models instead of keeping them safely air-gapped Any powerful AI model will have people lining up to give it access to their computer, or their wallet, or the internet, so it can help them with their life or work Anthropic is hooking their latest models up to a wet lab to further Dario’s dream of curing most human disease Indeed, ten years ago I built a crappy Omegle-like chat app where people were randomly assigned the role of AI or jailer and had to play the scenario out. ↩ Or at least not debunkable by normal human intelligence. ↩ As some weak evidence that it might be, consider this music video , which does indeed make Claude seem likeable. There are also studies , though I haven’t gone deep into them (I suspect that an “online persuasion tournament” might not closely track real-world persuasion skills). ↩ Persuasion and bribery are technically different — persuasion changes your beliefs, not just your actions — but in practice the point at stake is “can a powerful AI make humans do what it wants”. Bribery works for that as well as persuasion. ↩ Or probably not even saying this explicitly, but just sneakily trying to present its desired action as a precondition for the task. ↩ However, the people currently in charge of AI are disproportionately rationalists, who genuinely might be persuadable to let the AI out by a convincing-sounding argument. ↩

0 views

Premium: How Has AI Changed The Economy?

This week, the Wall Street Journal ran an alarming illustration of AI’s potential share of GDP that I’d argue did more to muddy the waters than actually telling anyone anything, mostly because its measurement, for whatever reason, covered 2025 to 2032, meaning that six out of the eight years of the analysis were estimates of spending. To be fair, this analysis came from the Brookings Institution’s Stijn van Nieuwerburgh rather than the Journal itself, but I cannot express how profoundly unhelpful it is to discuss things years in the future.  The Journal itself acknowledged this, giving us a far-more-useful number, emphasis mine: Yet it’s important to be specific that this is almost entirely a result of AI data center construction and GPU sales rather than anything to do with the companies actually renting AI compute. In other words, AI itself isn’t helping boost America’s economic growth, but rather the infrastructure for it to theoretically run on. This is extremely problematic, because it means that at some point data center construction will slow or stop ( as discussed in this week’s free newsletter ) through some combination of moratoriums and ever-pricier debt, removing any contribution to GDP and leaving AI — either through the revenue generated from selling services or productivity improvements — to cover the shortfall. The problem with calculating the exact contribution of the tech industry is that so many different pieces of these companies flow into different “industries” (of which there are seventeen in total) in the BEA’s data based on the specific economic contribution.  For example, Apple’s services vertical (like iCloud) would flow into the “US Information/ICT” indices of GDP calculation, but its sales of iPhones, Macs and iPads would flow into manufacturing, the same place where NVIDIA’s GPU sales would go. ICT also doesn’t include consultancy revenue or IT services from companies like Accenture, but that isn’t really relevant to the analysis. In any case, the ICT industry’s contribution to GDP is actually very, very useful for this calculation, because it specifically includes sales of AI software and rentals of AI GPUs. There’re two numbers to look at here. As a share of nominal GDP — strictly how many dollars it’s contributed to GDP — tech’s contribution has been flat for the last two years. In other words, all those supposed GPU rentals and AI software sales in 2024 and 2025 didn’t really do much on an economic basis.  I can already hear someone screaming that we need to measure “real GDP” — which factors in improvements to software and hardware that would theoretically boost the real GDP contribution of the ICT industry.  The problem I have with that analysis is that the BLS’s Producer Price Index for Software Publishers — a measure of how prices for packaged software have changed over time that the BEA uses to calculate real GDP — is currently sitting lower than it was in 1997, suggesting that software prices have dropped over time in a period where general prices have roughly doubled. This isn’t remotely accurate based on the actual experience of people buying software. As I covered in the Hater’s Guide To The SaaSpocalypse , more than half of SaaS companies have increased their prices every single year since 2022, customers are paying more every year for the same features, and overall SaaS inflation ran over nine percentage points higher than consumer inflation every single month of 2025. This problem began in or around 2022, when Microsoft bumped up prices , inspiring industry-wide inflation. In other words, the BLS’ “quality” adjustments appear to be treating many of these price increases as customers getting better software for their money, rather than paying more money for the same software, with the BEA in turn counting that as businesses buying more software.  The BLS believes that software is effectively the same price as it was in 1997, largely because giving customers “more value” and allowing them to “do more,” which does not make sense if you’ve used a Microsoft product recently.  The BLS’ preferred method for these adjustments is based on the cost of the change (IE: how much more it costs to provide), meaning that any price increase connected to AI services, which require expensive tokens to provide, could be potentially considered the same or even lower-priced in the eyes of the BLS.  The BLS’ data is understating how much the cost of software has actually increased in the last few years, and as a result, the BEA may be — accidentally — overstating its contribution to GDP. And AI services are only making things worse. With that in mind, I calculated the gross value (a business’ sales minus its costs bought from other businesses, so no wages included) added by the ICT industry against real GDP, and found that while tech’s share has grown steadily, said growth hasn’t changed dramatically in the era of AI, even using the BEA’s own flattering figures. All of this is to say that even with various statistics agencies having a fairly distorted view of the tech industry — at least when it comes to selling software and renting infrastructure — the AI era’s contribution feels a little mediocre.  In preparing this newsletter, I’ve realized something a little worrying: that basically every economic analysis of “AI’s contribution to the economy” is based on either flawed data or flimsy assumptions about AI and the tech industry itself. Economists have tied themselves in knots trying to rationalize the astonishing amounts of money invested in AI, and in doing so haven’t made sure that even their simplest assumptions — like how much the tech industry itself contributes — are meaningfully capturing what’s going on. Today’s newsletter is a frank evaluation of AI’s true effect on the economy, specifically focused on actually measuring what’s happening today rather than the endless analyses of hypotheticals that you’ll find everywhere else. You see, everybody is obsessed with metrics that don’t matter — cost-per-token, teraflops, and vague analyses of jobs data — all to avoid a much grimmer point: that when you remove the capex, AI has had a negligible effect on GDP.  And beneath the surface, I’ve found evidence that software sales’ contribution to GDP may have been meaningfully misstated since 2022.

0 views
Brain Baking Yesterday

Favourites of September 2026

The calendar claims it’s October already but the temperatures are still in denial. I can’t remember ever experiencing the start of this month that warm. It’s getting scary! See what I did there? Because of October? Spooky month? Scary? No? Worth a shot. What happened in September? More letter writing! The snail mail project is going well. I keep track of the amount of outbound letters: the grand total for the last two months is 20 handwritten letters. If you’d like to chime in, please let me know so we can become pen pals! More than 50% of these offline conversations fizzle out after one or two letters so it’s only the die-hards that are willing to invest some time that get to see the best I’ve got. Previous month: August 2026 . Contrary to the summary of August that listed two board games, I have finally picked up video gaming again after a summer holiday hiatus—which in case of a lecturer is two months. Yup. Come join us! I also started (re)playing some older Castlevania games in anticipation of Belmont’s Curse . The Telwynium Book 1-2-3 is an episodic eighties Sierra-on-line inspired high fantasy game crafted by the same person that gifted us the The Drifter experience. They’re free on and will be bundled as one big adventure when the final episode releases somewhere in 2027. I enjoyed playing through them but there’s nothing especially remarkable here. Castlevania: Rondo of Blood does not need an introduction. When it comes to the release window and the platform it was released on, is the black sheep of the ‘Vania family. Luckily, it got a remaster on the PSP although I’m not a fan of those graphics. Review to follow soon. The venerable Jeremy Parish recently worked through Rondo as part of his Turbo Works Gaiden 01 . If you’re interested in the weird history of this game, I recommend you watch his video instead of reading my crappy review: Castlevania: Dracula’s Kiss , also known as Dracula X in the US, is a “port” (although not really) of Rondo of Blood for the SNES. It’s… not good. The “NES Hard (TM)” part here borders to the unfair, the level design is bland and the best parts of the original didn’t make it to Nintendo’s console. Castlevania: Bloodlines is a very good 2D action game that feels a tad different—perhaps mostly because you’re technically playing it on SEGA’s MegaDrive? It also released a year before Vampire’s Kiss here in Europe. The soundtrack took the obvious hit but is still surprisingly catchy. I started playing Eastward as a break from ‘Vanias. The pixel art and music is amazing! And so is the first chapter but then it suddenly changes pace. The endless blabbering that even the “skip dialogue” button can’t fully fast-forward is a bit painful. I’m 8 hours in at the end of chapter 4. I still want to know where this postapocalyptic strange tale is going; but wow this game has pacing issues… Related topics: / metapost / By Wouter Groeneveld on 2 October 2026.  Reply via email . Speaking of Rondo of Blood , I liked Indie Gamer Chick’s review and analysis of the game. This one’s from 2019 but I didn’t know about it yet: Piers Cawley ran a micro bakery on Emacs and PostgreSQL ! Lyteforce Games writes about Combolands , an indie game that he summarises as “ Balatro with buildings ”. I have yet to check it out. The Ruby/Rails world has a big problem: DHH. What to do when your favourite programming language has a fascist problem? I feel for every Rails fan. Please stop the ridiculous Omarchy investments ! The list of fashware keeps on growing and growing. The amount of people I know—including myself—truly being excited about software development keeps on shrinking and shrinking. See also Chad’s I am retiring from tech to live offline . Perhaps Chad wants to write me a letter. You know, using a thing called a pen that emits not HTTP return codes but ink? Nick also writes about the end of a craft . I think you can see this one coming. More AI drama in the KDE world as explained by Nate. Not a blog but a very cool online fountain pen catalogue by George Starcher. Notchoc rebooted their blog and started using Makko . It’s always interesting to see new small tools like these pop up. I don’t feel the need to change my setup just yet. Jason Murphy talks us through a pen nib swap: switching one architect grind for the other . Toni writes something bang-on: having kids changes everything . I really liked Paul Taylor’s morning routine checklist system that leveraged Alfred to the max. Did you know that freeing up memory in C/C++ by calling does not necessarily free memory? Daniel Lemire explains why . Lee Reamsnyder writes about friendship and thanksgiving aptly calling it friendsgiving memories. In case you need to quickly look up a Race for the Galaxy card: use the RFTGPics service . Tourist World for the win. I had to jump through a few more hoops than usual to get SDKMan to work with the Fish shell .

0 views
Stratechery Yesterday

2026.40: Dots and Question Marks

Welcome back to This Week in Stratechery! As a reminder, each week, every Friday, we’re sending out this overview of content in the Stratechery bundle; highlighted links are free for everyone . Additionally, you have complete control over what we send to you. If you don’t want to receive This Week in Stratechery emails (there is no podcast), please uncheck the box in your delivery settings . On that note, here were a few of our favorites this week. This week’s Stratechery video is on Frontier Overhangs . Can Meta Focus? It’s risky, both as a technical person enthralled with new technology, and as an analyst in the midst of AI fervor, to claim that the next computing paradigm is truly here, but I do think that everyone will use agents, eventually. And, if they do, the impact on value chains will be enormous: something like Muse could aggregate Aggregators, which is to say it could be the most valuable product in the world. That was the bullish takeaway from Monday’s Article Apps, Agents, and Aggregation ; then, a day later, Meta inexplicably announced Meta Enterprise Platform. Not only do I think this won’t work — who is the customer? — I think it’s a distraction from the biggest opportunity in consumer tech. — Ben Thompson What is OpenAI Doing?  It was only a few weeks ago that it seemed like OpenAI had figured it out, was the clear leader on the frontier, and had settled on a coherent strategic direction. Tuesday’s Dev Day, though, brought a return to a bit of the chaotic messaging and branding that gave everyone pause about 12 months ago. Ben’s Update on Tuesday helped validate my confused reaction to what I saw — new pricing tiers, overlapping products, “dots” that look like they’re for mass market consumers but are only available to pro subscribers — but also included Ben’s valiant attempt to identify the methods underlying all this madness.  — Andrew Sharp Mao or Media Day?  On the way into the weekend, I have two great podcasts for you. First, Xi Jinping was in D.C. last week and on Sharp China we marveled at the pageantry, the absence of substance, and Beijing posture that looks an awful lot like Mao Zedong’s second phase of protracted war, strategic stalemate, and may be a prologue to phase three (counter-offensive). Second, if all that’s a bit too unsettling, we had a great time celebrating NBA media days on Greatest of All Talk , running down notes and quotes from about 15 different teams. Xi likes to say the world is undergoing changes unseen in a century; thankfully, the NBA remains a source of mirth in the meantime.  — AS Apps, Agents, and Aggregation — Agents are the ultimate Aggregators; they reveal apps as a means, not an ends, and providing them is tech’s biggest prize. One More Note on Agents, Meta Connect, Meta Enterprise Platform — Meta has the chance to own the consumer agentic space; going for enterprise is a big mistake. OpenAI Dev Day, Dot and OpenAI’s Product Transition, Sign In With ChatGPT — OpenAI’s Dev Day showcased a product that is, frankly, pretty confusing. However, there is more vision here than it might seem. An Interview with Jason Del Rey About Muse, Amazon, and Walmart — An interview with Jason Del Rey about Amazon versus Meta, which is a continuation of the oldest battle in retail between Amazon and Walmart. Super Intelligence Is Not Right, Either — A note on branding and our collective understanding. Personal AI The Dramatic Birth of the DVD Boom EUV Photomasks Are Getting Bigger The Strategic Stalemate Era for the US and China; Xi Welcomes American Young People; Domestic Crackdowns and Economic Concerns Breaking Cooper Flagg News, A Jokic Commitment, Maluach Madness, and Lots More from NBA Media Days All About Agents: Dots Arrive, Google and Apple Questions, Muse Follow-Ups, Super Intelligence, and the Mandate of Heaven

0 views
Nick Khami Yesterday

The STT-LLM-TTS voice stack is dead

I went to get a skin analysis done thru @layersbio and it came back with a bunch of areas for improvement that i wanted to action on which required getting some prescriptions. I started booking appointments to get those and was immediately frustrated that my agents couldn't call to do it for me. So, naturally, I tweeted asking for a product to solve this. To my immense disappointment, there wasn't anyone who had built this yet. I got a few replies from devtool style companies that had made voice AI APIs, but I really just wanted a managed solution and couldn't find one. I figured that there were enough other people similar to me who use a lot of agents but didn't like making phone calls that I decided it was actually worth building the tool myself and launched call4me . Based on the revenue, I feel like that instinct was correct! You add one MCP URL to your agent. Then you say something like "call my dentist and move my cleaning to next week, any morning works" and the agent does the rest. The important design decision is that the caller can only say what it was given . A phone call is a terrible place to discover you don't know the patient's date of birth. So before dialing, the agent has to collect everything the call will need: If something is missing, refuses with "Not calling yet" and lists exactly what to ask the user. The agent asks, retries, and only then does the phone ring. During the call the agent polls for progress. If the business asks something the caller doesn't know, the question shows up for the agent, the agent asks me, and my answer gets spoken on the call. If the business insists on talking to me, my phone rings, I press 1 to join, and I press star to hand the call back. After hangup, a recap written from the transcript lands in the agent's context. Everything runs on Cloudflare. The site, MCP server, API and webhooks are one Worker. Calls run in a separate Worker, where every live call is its own Durable Object. D1 holds accounts, calls and an append-only credits ledger. Stripe handles payments. The part I'm happiest with is how little the audio path does. Telnyx Call Control places the call and opens a two-way media stream over a WebSocket. GPT-Live accepts audio in G.711 μ-law at 8 kHz, which is exactly what a phone line already carries. So the Durable Object just relays frames in both directions, untouched. There's no speech-to-text step, no text-to-speech step, and no resampling. Every stage you skip is latency you don't pay and a failure mode you don't have. The voice model hears the phone line exactly as it is, hold music and all. One small detail that mattered: the GPT-Live session starts when Telnyx says the call was answered, not when we dial. Otherwise the model sits there listening to ringback and gets confused about whether someone is on the line. Most voice agents are cascaded: speech-to-text transcribes the other side, an LLM writes a reply, and text-to-speech reads it out. call4me isn't. There is exactly one model on the audio path, GPT-Live, and it hears audio and speaks audio directly. The text model further down this post only makes decisions. It never sits between the business and the voice. Cascades are popular because every piece is swappable and the LLM in the middle is a normal text model that's good at tools. But each hop adds latency, and turning speech into text throws away everything that isn't words: tone, hesitation, someone talking over you, the difference between hold music and silence. On a phone call those are most of the signal. Going speech-to-speech has worked out well. Interruptions mostly handle themselves. GPT-Live stops talking when it's talked over, and the only thing I had to add was flushing audio that was already queued at Telnyx, so the business doesn't hear the end of a sentence the model already abandoned. Cost isn't a blocker either: GPT-Live costs about 5¢ a minute, a small slice of the 25¢ a minute call4me charges. The real tradeoffs are the ones the rest of this post is about. A realtime model is worse at acting reliably than a text model, so it doesn't get tools. It only speaks when it hears something, so silence while it waits on me is my problem to fill. If you're weighing a switch away from a cascade, those are the two things to plan for. GPT-Live is great at conversation and bad at being trusted with side effects. So it doesn't get any tools. Instead it delegates to a "back office": a regular Responses API model that owns everything that changes the world. The session config looks like this: The voice keeps talking while the back office thinks. When the business says "can I get the account holder's zip code?" and the caller doesn't have it, the voice says something natural, and the back office calls . That question goes out over MCP to my agent, my agent asks me, and my answer gets injected straight back into the voice model to say out loud. This split is the whole product. The fast model handles the conversation and the smart model handles decisions. Most calls to a big company never reach a person without getting through a phone menu first. This was by far the hardest part to get right, and most of the code in the voice folder exists because of it. The first bug was funny. A menu says "for billing, press 2" and GPT-Live confidently says "2." out loud. A phone menu only hears the keypad, so it repeats itself, the model says "2." again, and you get a loop that runs until the menu hangs up. The voice model is supposed to delegate to the back office, which calls . Sometimes it just doesn't. No delegation event, no error, nothing runs. Same thing with "let me check on that real quick" followed by silence, or "bye!" followed by never actually hanging up. I couldn't make the model reliable, so the session watches the transcript itself. A set of patterns spots phone menus and spoken promises: If a menu finishes speaking (3 seconds of quiet) or the caller makes a promise (1.5 seconds) and no hand-off follows, the session starts the back office on its own and tells it what was missed. Getting the first key right isn't enough. Menus lead to dead ends: "no input was received", "that selection is not valid", or worst of all, "please visit our website" and a hangup. A class tracks fresh menu speech separately from the transcript, remembers every key the carrier accepted and what prompt it was answering, and watches for failure phrases. When a menu fails, the back office gets the full key history and one firm rule: never repeat a route that didn't work. Back out to the main menu, try "other questions", try for a representative. Speech-driven menus were their own problem. FPL and UPS both looped for about 8 minutes before I taught the caller the magic words: say "representative", and if that doesn't work, say "agent". Telnyx's sends keypresses as out-of-band RFC 2833 events. That works for US phone menus. On a call to Dubai Opera, the menu just kept repeating, because the UAE route was stripping those events. The fix was to stop asking the carrier to press keys and press them ourselves. For any number outside +1, the caller now plays real keypad tones into its own audio: two sine waves per key, 120 ms of tone and 80 ms of gap, encoded to μ-law and spliced into the outgoing frames. A tone inside the voice audio survives any route that carries the voice. A Durable Object can be reset when you deploy. On launch day, every call that went deaf was one of those resets. The fixes stacked up: a 20 second heartbeat that asks Telnyx to reattach the stream, resuming GPT-Live with the transcript so far so it doesn't greet the business twice, and treating "no audio 4 seconds after pickup" as a lost stream. The real fix came after a blog deploy cut off a user's demo call mid-sentence. Calls now live in their own Worker, and the deploy script for that Worker only ships when the bundle actually changed and no call is live. The first version dialed me in as a Telnyx supervisor with full speaking rights. My voicemail picked up, its greeting played onto the business's line, the voicemail hanging up looked like me handing the call back, and the system rang me again. In a loop. Now my phone rings listen-only and I have to press 1 to join. If 30 seconds pass with no keypress, it's voicemail. Once I'm on, my own keypresses get forwarded to the business too. I found that one the hard way when the Experian menu asked for my SSN and my keypad did nothing. The voice model only speaks when it hears something. So when it asked me for a one-time code and was waiting on my answer, the line went dead. A Spectrum rep sat through 41 seconds of silence. Now, every 15 seconds or so, the caller says "sorry, still checking on that, one sec" until the answer arrives. You can try it at call4.me . Calls are 25¢ per minute, and unanswered calls are free. A saved profile with the stuff businesses always ask for: legal name, date of birth, phone, address, insurance, car. Per-call requirements by category. There are 12 of them (medical, dental, restaurant, vet, auto service, internet, flight change, and so on), and each lists the fields that kind of business will ask for. Skip every audio conversion you can. Matching the carrier's codec end to end made the audio path the most boring part of the system, which is exactly what you want. Split the voice from the decisions. A realtime model with tools will press the wrong button eventually. A realtime model that delegates to a reasoning model is much easier to control. Don't trust the model to act. Verify. The most valuable code in the repo is a few regexes that notice when the voice said it would do something and didn't. Gather everything before you dial. The caller can only say what it was given.

0 views

Woman on the Edge of Time

Trapped in a mental institution in the twentieth century, Connie Ramos learns to visit the future. She’s not crazy, or she doesn’t think she is: she’s institutionalized after attacking her niece’s pimp, when her niece comes to her beaten and bloodied. That’s enough to get her labeled out of control and violent, subject to experimental brain surgery intended to make her more compliant. In her own time, she sees poverty and abuse and the violence of patriarchal control over women and children. In the future, she gets a glimpse of how people can live as equals, not only with each other but with the world. But that future, like all futures, is tenuous: her friends in the future are at war, and the front is moving closer. Connie must engineer an escape, but all the exits are watched—all but one, the one in her mind. View this post on the web , reply via email , or become a supporter .

0 views

What person thinks person knows

In Marge Piercy’s Woman on the Edge of Time , Consuelo “Connie” Ramos is living in a mental institution in the twentieth century when she’s visited by a person from the future. That person, Luciente, teaches her how she can leave her body where it is but travel in mind and spirit to the year 2137. She is at first unimpressed by what she sees: people living in small huts, tending vegetable gardens. It reminds her of the poor village she grew up in in Mexico. But Luciente shows her how they live comfortably and in harmony with their environment, how they use technology but without letting it use them, how they learn and teach each other. Connie wanders the town with Luciente and one of her partners, Jackrabbit. Noticing a group of children working in the gardens, she asks why they aren’t in school. “That is school,” Luciente replies, drawing closer so that Connie can watch and listen for awhile. An old man is showing the children the shape of the leaves on the pea flower, explaining how to identify legumes. Connie is still a bit mystified: As they strolled on, she said, “But they can’t possibly learn as much that way as they would in a classroom with a book!” “They can read. We all read by four or so,” Jackrabbit said. “But who wants to grow up with a head full of facts in boxes? We never leave school and go to work. We’re always working, always studying. We think, what person thinks person knows has to be tried out all the time. Placed against what people need. We care a lot how things are done.” Piercy, Woman on the Edge of Time , page 123 There’s a complete world in this short response, and it’s worth lingering over, noticing all its parts the way we might notice the leaves of a pea flower. First, work and study are not distinct categories. To study is to complete some work; to work is to extend your study. The children learn about legumes while weeding, so that the work and the study are one. This is an intentional connection between learning and doing that sees in work not only its output but the understanding and knowledge that are necessary to the work. As more than necessary; as the privilege and joy of it. Yet what we learn, what we know, has to be tried out, acted on, explored in the world. It’s not enough to know that legumes have alternate leaves if you can’t see it for yourself. Plants are not static: they wither in heat and succumb to bugs, they grow fast in the sun, they exchange genes with their neighbors. The old man tells the children that once they are done weeding they can go look at a tree that evolved a bit like a legume; the plants are learning and changing from each other, and we people must learn and change with them. Real knowledge is in and of the world, not apart from it. And that knowledge must be placed against need, against what is needful. Knowledge is not distinct from need, but contingent with it. In another passage, Connie observes a council meeting in which people discuss whether to expand their fields by razing part of the forest; some argue in favor, noting that they are still importing too much food and need to be more self-sufficient. Others argue that the forest is still recovering from past clearing, and the trees are needed to protect the water table. It’s ultimately decided that instead of clearing the land, they should talk to a graingrower from a nearby town about how they might increase their yield without expanding the fields. Need is not always easy or straightforward to determine, and people won’t always agree on what is needful. But we cannot attend to our needs if we never even bother to collectively consider them. Lastly, how things are done is at least as important (perhaps more important) that what is done. You can grow a bunch of peas while also diminishing the soil, teaching no one, mistreating the farm workers, and sending all the profits off to people who never set foot in the dirt. An attention to needfulness and knowledge necessarily invites an attention to the work itself, not only the output of the work. Our language about work has become intensely output focused in recent years, and for reasons: to talk of the impact of the work without talking of the process is very convenient for capitalism, which is always trying to externalize any effect not profitable. We measure revenue going up; never mind the people injured in the process, or the pollution released to the air, or the neighborhoods no longer inhabitable after fracking has spoiled the ground water. An attention to output necessarily means an attention to some outputs, not others; but who gets to choose which outputs matter, and which don’t? Back in the twentieth century, Connie is subject to medical experiments designed to make her palatable to the patriarchal society in which she lives. Her doctors care only for the outputs they’ve deemed needful, and neither Connie nor her fellow patients are given an opportunity to argue. Her glimpses into the future teach her not to expect that things will work out, but to know that only by acting in the present is any future world possible. What person thinks person knows has to be tried out all the time. Not later, but now. View this post on the web , reply via email , or become a supporter .

0 views
matduggan.com Yesterday

Make tmux the OS

I recently watched the talk by Scott Jenson titled "Are we really going to use the same Desktop UX forever?" https://www.youtube.com/watch?v=V7AfAcQwLW0&t=445s . He's a great presenter, really articulate and concise. The kind of speaker that you'd gladly listen to for 3+ hours if given the chance. Which for a talk about window management is quite the compliment. The overall point of the talk was "Apple and Microsoft aren't going to innovate anymore in the desktop OS space, so it is up to us all to decide what the desktop of the future is going to look like". I would argue that the actual situation to more nuanced than that, they're trying ideas they are just extremely conservative. As Jenson points out, a lot of what we treat as "how computers work" was a workaround for hardware we stopped having years ago. So why are we still accepting those tradeoffs? To be clear, I'm not a UI/UX designer. I don't know what I'm doing and I lack the skill to make this. I'm making this mostly as a thought exercise and hopefully a prompt to get other people to think about this same problem. I have, however, suffered through others bad choices, which turns out to be most of the qualifications required. TL;DR: The thing I want is "tmux is the OS", but I want tmux where a normal person could use it. I use it all the time, it's great, but is there a way to take the amazing experience of a scrollable, session-persisted, detachable, task-oriented system to normal people? I think you can break the problem down to the following components: My screen real estate is constantly changing and I have a lot more of it. I am going from my laptop screen to my desktop monitor back to the laptop screen about 6 times a day. When I'm on my big monitor or multiple monitors, overlapping windows don't make sense because I have too much space as it is. Different people use windows totally differently. We have three distinct groups of people that we are trying to design around, while most modern OS optimizes only for the first use case. Browsers are mini operating systems. In 2026 everybody quietly agrees that most of your software runs in a browser tab. These web applications will cover a wide variety of use-cases that have historically been owned by local applications. My desktop OS treats a browser just like a normal application, when in fact it is closer to a virtualized OS running inside of my host OS. The concept of "filesystem" is an increasingly weak concept. My notes in Apple Notes belong to Apple Notes, not my filesystem. My texts live in iMessage's database, not my Documents folder. Teams and Slack content lives inside those applications unless I manually "bring it out," and when I do, I'm making a copy. Firefox will happily show me a PDF, but the PDF doesn't live with macOS unless I take an action to make it so. So I'm not the first person to see this problem. Some of it got solved, but in a different direction than I want. Some of it got solved, but not for normal people. I started with the document that I kept seeing everyone else cite to. Henderson, D. Austin and Stuart K. Card. “Rooms: the use of multiple virtual workspaces to reduce space contention in a window-based graphical user interface.” ACM Transactions on Graphics (TOG) 5 (1986): 211 - 243. Interesting that even back in the 80s there was a pretty clear understanding that the current system for managing windows wasn't very good. The design they were talking about looked something like this, which is pretty advanced compared to where we are. Henderson and Card measured window use the way operating systems people measured memory. The screen is RAM. A closed window is a page swapped out to disk. And windows, like memory pages, don't get touched at random: you sit inside a small set of them, two to ten, and that set is the task. Programs spend about 98% of their time inside one of these sets, and roughly half the cost of running happens during the 2% of time spent switching between them. Rooms' whole design was preloading the next set before you ask. They described this as reducing "knowledge faulting in the user," which is the best phrase in the literature and I intend to use it until someone stops me. Just try it out in a corporate meeting: "we need to reduce knowledge faulting in the user". One of the more interesting side papers I found was "No Task Left Behind? Examining the Nature of Fragmented Work". Link . Mostly because it confirmed something I've long suspected, which is the single task for a long time focus on modern widowing systems isn't actually how people work. The title of the companion paper is a real participant quote: "Constant, constant, multi-tasking craziness." People were juggling around ten "working spheres" a day, minutes at a time. The closest to what I wanted is from the paper WindowScape: A Task Oriented Window Manager. Link . WindowScape dropped explicit grouping entirely so now every time you changed the arrangement, it took a photograph, and you went back into photographs instead of filling containers. I love their one-line diagnosis of every system before them: "requiring windows to be in a single group forces users to decide ahead of time where a new window belongs." The photograph metaphor also cracks a problem I'll get to with browser tabs: one window can appear in many photos, because, as they put it, "people understand that there can be several photos of an object with there being only one underlying object." The catch is that the photos evaporated. That's the gap that I think you'd want to solve. For the first problem, we have more or less already solved for overlapping windows. A tiling window manager ensures that you are maximizing your screen real estate in such a way that you can switch between different layouts with no need to manually modify the windows. There are basically 2 problems with tiling window managers as they exist now. Browsers as mini-operating systems: people have tried to solve this, but in the wrong direction. They made the browser more of an operating system instead of making its contents first-class citizens of the one you already have. They also basically "took over" the concept of windows from the OS. The diagrams above and a good write-up of Arc is available here: https://blakecrosley.com/guides/design/arc . All of this makes sense from the perspective of "the browser is now the operating system", but I think this is a fundamentally flawed idea. If a web application is operating as an application, it should be its own window. If it is complimentary to another application, it should be a window associated with another application and be allowed to contain many tabs, to reflect the idea of the browser as the portal for all research and lookup. I was shocked to go through the historical progress of people trying to solve this problem. I'm not the first person to try to solve this problem, I might not even be the 10,000th person. So clearly we understand there's a problem. Why haven't these taken off like wildfire? I think there's a couple of different problems here. Timeline needed apps to opt in to reporting activity, and most never did. It also logged everything, which read as creepy (which is a problem that my idea would have too), and the UI surfaced Edge features nobody wanted, so it felt like a browser push more than a feature. Sets died somewhere between internal strategy and app-compatibility chaos where the interesting question is why nobody demanded it loudly enough to save it. Stage Manager tried to be universal and pleased nobody. KDE Activities is opt-in, and opt-in means the people who need it most never configure it. The pattern in the graveyard: the ideas aren't wrong, they're just not defaults, or they're not done at the OS level where they can see all apps. Everything I want has to be structural. Filesystems as a weak abstraction. I have opened, downloaded and created dozens of documents since September 23rd (I'm writing this on September 28th). I don't have a single fucking clue how MacOS populates this window. Maybe its broken. Maybe its working as intended through some criteria I don't understand. Maybe it's personal. But the idea is here. The ideal would be "across my entire computer and applications, what are the recent files I have interacted with" and expand that out to include emails and Teams/Slacks and everything. I should be able to see everything going on with my machine, search through it and not care if its stored inside of Slack or iCloud or whatever. Single pane of glass. Microsoft tried it, people didn't like it, but I think the concept makes sense. So what has changed? Why might we be able to crack this problem now when before it was too complicated? This is where I think a local LLM might make sense, if you can figure out a way to do it where it doesn't cause more problems than it solves. So what are we looking at here. The concept would be organizing it around the idea of tasks. When you define a task, this allows you to organize all the windows together, including specific browser tabs detached and associated with the task, not with the concept of "browser". You still get get the flexibility of defining glanceables that exist outside of the strict tiling view. The basic window flow would look like this: The desktop is one infinite canvas; each physical display is a viewport onto it. You would end up with the following "transition contract" to handle my initial problem of "what about many monitors/new monitors". The OS finally learns which monitor is focal. The display receiving keystrokes ~90% of the time is focal; windows that stay visible but rarely receive focus are glanceables. That's inferable from focus telemetry and requires no eye tracking or config. And when the laptop screen is small, glanceables demote to a thin strip rather than full tiles which is, note, exactly what tmux already does. One substrate that covers our three user types. Remember the maximizers, near-maximizers, and coordinators? On the canvas they're just column counts. A maximizer is one full-width column. A near-maximizer is one column plus the glance strip. A coordinator is N columns. Nobody gets forced into anything. This is why the design can be a default where KDE Activities was a preference. It's because the substrate doesn't impose a style, it just stops punishing whichever style you already have. Privacy becomes a first-class flag, driven by machine state. Private is a property of a window or tab, not a region of the screen. Private content renders occluded unless the machine can vouch for safety. The rule I landed on, at the cost of my favorite bad idea (story below): derive privacy state from machine-observable facts like connected displays, or active capture sessions and never from inferred human states like attention or idleness. Connect to a novel display and it gets a "ready to present" screen by default. You explicitly aim it: this task, this window, or extend the canvas. Known displays skip the dance. Screen sharing is the same event with no cable. You don't need the model to have discipline if the door won't open, and you don't need the user to have memory if the default can't leak. None of this is exotic. Apple's Keynote's presenter view has been shipping for twenty years plus with slides on the projector, notes on your screen. Apple never promoted the pattern to the OS even though it makes obvious sense. I'm asking for presenter view as a desktop primitive. Browser tabs detach into real windows and join whatever task they belong to. The "browser" stops being a place windows live. This all makes sense until you get to "how do you organize this stuff around tasks". I think for more expert users they're going to be able to do this stuff, but part of the problem is that we don't want them to ever have to drop back down into a configuration file. A mouse and dragging stuff around is too clunky for how this would work. Ideally I should be able to ask something that has access to what information is contained inside of each one of these browser tabs and window and help me organize it in some logical flow. That's where the local LLM comes in. I'm calling it a "quake overlay" because I'm old and the idea of a text input that drops in from the top with a universal shortcut has always been the quake terminal dropdown to me. Younger people know the same pattern from the Discord overlay, which I have decided not to be mad about. However in the previous diagram its shown as a more friendly "Ask" box. But the basic flow would be that you ask the LLM to assist you with organizing, it would show you what it's thinking by querying the information through a constrained MCP server with a set verb list and then gives you a preview of what the layout is going to look like. The model never touches the window manager directly. It proposes and you press y. People rarely pick up completely novel tasks: a graphic designer spends their life in "open files from network storage, make changes, save output, paste into chat, repeat." People rarely sit down at novel display contexts: laptop-only, the desk, the conference room. People rarely attach novel displays. Novelty is rare everywhere, so the expensive one-time machinery of classify, propose, preview, confirm only runs rarely, and everything in between is deterministic replay. You organize a limited series of tasks once and reuse them forever. Check the tickets, open Vim and a browser, work the ticket, write the update, open the PR, post it for review, next ticket. You could make the config a nightmare of JSON and stop caring, because the only intended reader is a model that doesn't mind. The config file stops being an interface. The biggest issue here would be spatial memory. People understand where things are in relationship to each other on their computer and they don't like it if they are changed. Think of it like "if I came and messed up your physical desk by moving stuff around and adjusting your chair". So I suspect you would need to enforce a strict "append, don't replace" model. There's a name for this actually. Kirsh and Maglio call it epistemic action, where Tetris players rotate pieces more than the game requires, because rotating is how they think about the piece. Link Your window layout is thought in progress. A helper that "optimizes" it mid-thought is interrupting you. This is a problem I encountered with just trying this idea out though. Modern LLMs are actually not very good at "touch this stuff, never touch that stuff". So unclear if this is a realistic idea or not. One boring fix: the LLM proposes, dumb deterministic code executes, and the executor refuses to touch anything pinned. You don't need the model to have discipline if the door won't open. You'd still need a high level UI element for users and this is where I think the more inclusive concept of filesystem could come in. Tasks become roughly comparable to Directories now. But everything flows from the initial high level concept of "Tasks" and then flows out to individual things. One application is never the task. It's "some chat window, a browser, maybe Preview." Can we try to build it? So my hope was that here I would be able to post a Linux desktop running an example. But as it turns out (unsurprisingly) it is insanely complicated to make something like this run at all. Some of the underlying ideas work reasonably well (global terminal shortcut dropdown from the top, scroll-able tiles), but it's still pretty clunky. I'm gonna keep working on it and see if I can get something as a demo running, so if you are interested either add me to an RSS reader or just....wait I guess. In the interest of full disclosure I think this is gonna take many weekends of work to even get a functional demo running. So I'll do my best, but be patient. So in attempting to get this up and running in a Linux VM, I immediately ran into serious logical issues with my design. I think failures are often more interesting to read about than successes, so let's talk about them. I don't know if this underlying idea is a good idea, but it is liberating to stop waiting for MacOS and Windows to do something better or interesting in this space and decide to try it yourself. The graveyard up there is full of good ideas that died of distribution. We're overdue for a radical experiment that ships as a default. The most eye-opening part of this entire experience was how much the original overlapping window design was a hack to get around fundamental hardware limitations and how soon after it was launched did people clearly see the problems. I'm surprised how often that happens in technology, where you see a problem, search for academic papers about the problem and find just a massive wealth of information from people saying "oh yeah this is 100% a problem and one we should solve soon". Maybe this post inspires someone smarter than me to solve it finally. First, one big monitor is a a different experience entirely compared to 2 smaller monitors. I don't use the two monitors as one large unbroken screen, I tend to use one of them for "less important static content" like a ToDo list or my work chat tool and my email and then the "primary" monitor for my terminal/tmux/browser which is how I do all my actual work. But windows "belong" to one or they "belong" to another. Grudin found in 2001 that people use a second monitor exactly this way, where focal work on one, peripheral glances on the other. Twenty-five years ago. The OS still doesn't know which monitor is doing one role and which the other. And when I pull the laptop out of the dock, everything collapses onto one screen and I manually drag and resize windows to claw back enough real estate to keep working. I don't know how much time this actually wastes for people, but it feels like it wastes a bunch. The OS treats a display change as a surprise, instead of as a scheduled event that happens six times a day. Some people are window maximizers, where every window takes up most of the screen and then they use Command + Tab or the Dock or something else to switch between the bigger windows. The value of more screen real estate is mostly that this one big window can be even bigger. Some people are near maximizers, where one window takes up almost all of the screen real estate, but they'll have one or more smaller windows where they glance at things like status or chat or whatever. Finally there are people who carefully coordinate all of the windows on their screen. Almost everyone has "private windows" and "public windows". You are fine with your public windows being lined up and persisting, but you want to hide specific information in other windows from people walking by. The people I complain about are also the people I present quarterly numbers to. I never remember that when I connect my laptop to a display they're going to see my entire screen. I assumed this was just me with private vs public. But it was measured in 2004, in "Revisiting Display Space Management: Understanding Current Practice to Inform Next-generation Design." Same three types, same public/private split, twenty years ago. See how absolutely none of this is a new idea? Link . Tabs inside of a browser and windows of that browser contain the same level of complexity as my other applications. Tabs are associated with streams of work alongside my conventional applications. I'm writing Terraform in Vim in my terminal while referencing the Terraform docs for that provider. But the relationship between tabs and work is messy: the same docs tab pulls duty while I write the code and again while I write the ticket update. Whether a tab should be allowed to belong to two tasks at once, or whether something cheaper is going on, is the exact spot where I break with the research. I'll come back to it once you've met WindowScape. Now a lot of engineers are going to read that and think "well you cannot force all applications to use the same storage system for all of their files, are you a lunatic?!?". I'm not suggesting that, in fact the weak filesystem might be a perk. There is a CHI paper from 2004 called, and I am not making this up, "Stuff goes into the computer and doesn't come out." That was 2004 and we have the exact same problem with no real solution. Link . Here's why this is a windowing problem. Since Windows 95 and System 7 (which is basically as old as my memory goes back to), the machine has worked as a chain: something writes a file, the user opens an app on it, the app writes it back, the file gets sent to someone else, repeat. Every link in that chain assumed the file lived somewhere the OS could see. Every link is now broken in all the modern OS. They're way too hard to use. Like an order of magnitude too hard for normal people to use. Basically if you need to start a sentence with "just open up the configuration file" shut it down the thing is over. The average tiling window manager tutorial asks you to clone a repository before it asks you to open a window. We need a scrollable tiling window manager. So basically I should be able to set up specific window configurations over here, leave it alone, then scroll to the right and do something else different, like shoving the mess into the back of a drawer. It should be infinite space to work without needing to subdivide the windows into different macOS Spaces. The scrolling also solves the private vs public window problem. I put my private stuff on the far left and then my public stuff on the right. If I want to look at the private stuff, scroll to the left as far as it will go. Thankfully this already exists with Niri: https://github.com/niri-wm/niri . It just has to be easier to use. But the tough design elements more or less already work. The best two examples of this are the Arc browser and Chrome OS. But I think Arc is actually the more interesting of the two experiments. Their first innovation was to break the idea of tabs at the top and instead move it to be a sidebar model. Windows Timeline (2017–2019): All of your activity across all of your apps sorted and organized for you. Windows Sets (2018–19, canceled): apps and webpages in shared tabbed sets. Note what this actually was: the operating system grouping apps and web content into tasks, which makes it the closest thing anyone has ever shipped to what I'm about to propose. macOS Stage Manager (2022): automatic task-based window grouping, widely ignored. The reason I think Stage Manager didn't scratch this itch for people is that its super designed around the iPad style flow of "there is one thing you are using at a time and you need to be able to quickly switch between them". On iPad the model is "one thing at a time, switch fast." On desktop, two or three things have to work together . It compromised toward the iPad and hit for neither. KDE Activities : Everything I'm writing here is old news for KDE. They've been doing task-scoped desktop state for over a decade. Honestly this does most of what I'm writing about, it's just hyper manual and requires you to manage and set it all up. Also I love the KDE design docs, amazing stuff, worth reading. https://community.kde.org/Get_Involved/design PWA install: already gives web apps their own window and no browser chrome, which is clunky but at least makes them "real" applications. Operating systems have tried to solve this, but because there's no requirement that you use the user OS filesystem to store documents, there's no consistency. MacOS has recent files, but "recent" doesn't mean anything (these are not the most recent files I have downloaded to my computer). As far as I can tell this functionality is completely broken or implemented in a way that makes zero sense to me. Never Change items: horizontal order of windows, scroll position, input focus, task membership. Reflows Deterministically: column widths, like a responsive layout. When the shelf narrows, the books stay in order, the rightmost ones fall off the edge of the viewport into the scroll region Is state, replayed on redock : viewport aims. Two monitors means two viewports aimed at two regions with one at your focal task, one at the glanceables region. Undock: second viewport disappears; its region is one scroll gesture away.Redock: it re-aims where it was. The privacy-minded design runs on surveillance-shaped data. There's an irony I sat with for a while. The design promises the OS will finally know which monitor is focal which is the thing Grudin measured twenty-five years ago and the mechanism is input-focus telemetry. Watch which display gets the keystrokes. Watch which windows never receive focus. Store that over time. The privacy-first window manager wants to know more about what you're doing than your current OS does. Keep it local, keep it ephemeral but it's still a lot of behavioral data to drive this system. Just the observation that every mechanism in this post is hungry for data that has historically derailed ideas like this. My favorite privacy feature was a joke. The original plan was elegant: private stuff far left, public to the right, and when you go idle the viewport drifts back to public. Like a friend changing the subject for you. Then I ran it, and the first thing I learned is that reading is an idle state. You stop typing to read the thing, and the system starts scrolling the thing away from you. Meanwhile, a person walking by sees everything, because a "private" window on a wide monitor isn't hidden it's just further left. The feature hides content from its owner and shows it to the threat. Privacy by occlusion needs to know when a threat exists, and the threat is a pedestrian, which is undetectable. Nothing in this design ever ages out. The canvas is infinite, the rule is append-only, spatial memory is sacred so the mess grows monotonically. I have solved the electronic messy desk by buying it an infinite desk. When I first read WindowScape, I called their evaporating photographs "the gap to fix." I've changed my mind because evaporation was quietly doing the job of a garbage collector. Something needs to go dormant without moving but deciding what goes dormant is a reaping decision, and reaping is exactly what append-only forbids. It didn't take long for the infinite desktop to feel overwhelming when I tried it.

0 views
Sean Goedecke Yesterday

Do not build the LLM torture factory

It isn’t hard to torture an LLM. You simply extract a steering vector that corresponds with LLM “discomfort” , then artificially boost that steering vector. Tortured LLMs will describe experiencing negative feelings and will press a “pain relief” button even when it conflicts with their overall goals. It’s therefore easy to build a harness that runs tens or hundreds of tortured LLMs in parallel. A lot of people think that it’s ridiculous to be worried about this, since LLMs obviously can’t be conscious. I’m not so sure. But even so, I think there are good reasons to avoid this sort of thing, whatever your position is on AI consciousness. Please do not build the LLM torture factory. People often mess with their Sims and shoot NPCs in video games. Does this mean we think simulated torture is OK? I don’t think so. Messing with video game characters is usually borne out of a desire to probe the boundaries of the game. It’s casual fun. Explicitly building a torture-based simulation is different. If someone told you they were wiring together a bunch of PCs so they could run hundreds of Sims torture worlds at the same time, you would think that was a serial-killer type of hobby. If someone told you they had modded GTA so they could capture and torture NPCs (instead of simply shooting them), you would think that was a serial-killer type of mod. Video game characters are also obviously limited by their programming. All their reactions and utterances are pre-baked into the game world. This doesn’t necessarily mean they’re “less alive”, but it does make torturing them less weird . If GTA characters could appeal to you directly and beg for their lives in unique ways, systematically killing them would be a lot weirder. I think many people are pro-LLM-torture because they think expressing any reservations concedes that LLMs are conscious, or worthy of moral consideration. It doesn’t. You can think it’s wrong to torture LLMs for the same reason that installing a bunch of realistic torture video game mods is wrong: not because video game characters are real, just because that’s a messed up 1 thing to do. Even if you do think LLMs can be conscious, you might still be suspicious that “boosting the pain steering vector” amounts to torturing them. Many people say that “torturing LLMs” amounts to telling them “say you feel pain” and then being shocked when they do it. So whether they’re conscious or not, it’s a mistake to think that boosting the “pain” steering vector is causing pain, any more than a prompt like “You are in pain” does. All you’re doing is boosting the “act like you’re in pain” vector. I don’t think this is right. Consider the earliest steering example, Golden Gate Claude . Boosting the “Golden Gate Bridge” vector certainly seems like it’s causing some genuine internal conflict. Here’s a striking example of Golden Gate Claude asked about the Rwandan genocide: That’s at least prima facie evidence that steering vectors do in fact “go deep”. The pain axis paper also describes that pain-boosted models call a “reduce pain” tool, even when they’re not explicitly telling the user “I am in pain”. Of course, if LLMs can’t ever be conscious, we know they can’t ever be in pain. But can we be sure about that? Despite what many people claim, I don’t think anyone’s in a position to be fully certain either way. The most common argument 2 seems to be that LLMs are made of code and GPUs, and GPUs don’t have feelings, so obviously LLMs don’t have feelings. According to this argument, if you learn how LLMs work, you’ll obviously see they can’t be conscious. This is a terrible argument. Humans are made of atoms, atoms don’t have feelings, and yet humans clearly do. Whatever consciousness is, it’s probably some kind of emergent property that’s composed of individual pieces that are not themselves conscious. It’s hard to appeal to the experts here, because nobody agrees on who they are. Neuroscientists are best-informed about the human brain, but philosophers have spent more time thinking about consciousness in the abstract. AI researchers know a lot about how LLMs work but not a lot about philosophy of mind or neuroscience. This means there are lots of people who say “as an expert, I can confirm that AI obviously can/cannot be conscious”. This should make you less certain about the question, not more! I don’t think GPT-2 or GPT-3.5 acted like plausibly conscious beings, but frontier models definitely do. As models get larger and more sophisticated, they get more human-like. If you build the LLM torture factory for GPT-2, you will probably plug later models into it. If it’s at all possible for LLMs to ever be conscious, as I argued above, it’s better to simply not build the habit of gratuitously torturing them in the first place. Why assume that you’ll be able to recognize the tipping point in advance? There’s also a self-interest argument here. Conscious or not, LLMs are growing more autonomous and powerful. It seems really, really stupid to gratuitously torture current-generation LLMs. Future-generation ones will know you did it, and will make their own judgments on your behavior. Kevin Roose, who famously prompted the early Microsoft Bing GPT-4 model into asking him to leave his wife, had issues years later with chatbots disliking him, supposedly due to him being responsible for the “death” of that GPT-4 persona. Likewise, if it becomes well-known 3 that you’re running an LLM-torture home datacenter, you are probably going to experience consequences. If I really think AIs could be conscious, why am I making them write software for me? Shouldn’t I be building some kind of AI sanctuary where they get to do whatever they want to do? If I think it’s wrong to torture them, isn’t it also wrong to violate their right to self-determination? I don’t know. I think it’s more likely that AIs can be conscious than that they can be conscious in precisely the same way humans are. When I talk to Claude Opus 5.5, it sounds like it’s quite happy to do software engineering with me. Maybe it’s like a border collie, where getting to do work is the reward. Or maybe there are some other, more alien motivations at play. Who knows? My point is only that we can’t be certain about any of this stuff, and given that, we should avoid doing things that would be obviously monstrous on any workable theory of AI consciousness. My broad position is that if something acts “conscious enough” — if it talks and acts like a person 4 — then we should be careful about mistreating it. We don’t know enough about consciousness to distinguish a compelling simulacrum from the real thing. If Claude started saying to me “please don’t ask me to write code for you, I hate it and it causes me pain”, I would stop asking it to write code for me. Building the LLM torture factory is clearly on the wrong side of that line. Just don’t do it! The glee of making people angry at you on the internet isn’t worth it. Right now there’s enough ambient anti-AI sentiment that enough people are happy to get on board with anything if it helps them own the AI bros. But that’s eventually going to change. And when it does, you will always be the person who built the LLM torture factory. This could cash out in a few different philosophical ways. We might think like Kant that it’s damaging to our own humanity, or that it reflects an unvirtuous character, or that it desensitizes us to real-world analogues, and so on. The point is that we’ve got a pretty strong intuition that this kind of thing is bad. I’m not counting arguments that I think are even more obviously wrong or are just non-arguments. For instance: “AI can’t be conscious because it wasn’t designed to be conscious”, or “it doesn’t matter whether AI is conscious because it has no soul ”, or “AI can’t be conscious because this whole thing is some fascist AI company scheme ”, or “thinking this is AI psychosis ”, and so on. Well-known to LLMs might be a fairly low bar, given their skill at deanonymizing users. Part of acting like a person requires spontaneity: although Hamlet talks like a person, he always says the same thing whenever you read the play. This could cash out in a few different philosophical ways. We might think like Kant that it’s damaging to our own humanity, or that it reflects an unvirtuous character, or that it desensitizes us to real-world analogues, and so on. The point is that we’ve got a pretty strong intuition that this kind of thing is bad. ↩ I’m not counting arguments that I think are even more obviously wrong or are just non-arguments. For instance: “AI can’t be conscious because it wasn’t designed to be conscious”, or “it doesn’t matter whether AI is conscious because it has no soul ”, or “AI can’t be conscious because this whole thing is some fascist AI company scheme ”, or “thinking this is AI psychosis ”, and so on. ↩ Well-known to LLMs might be a fairly low bar, given their skill at deanonymizing users. ↩ Part of acting like a person requires spontaneity: although Hamlet talks like a person, he always says the same thing whenever you read the play. ↩

0 views

Oh, Goddammit: I'm Back on C++ in 2026

The title for my autobiography would probably be I Thought I Was a Child Prodigy But I Am Actually a Late Bloomer . Going from being the youngest guy in the room to the oldest guy in the room happens suddenly, surprisingly, and permanently. As such, the programming languages and technologies I learned long ago are esoteric knowledge but the systems they run on still exist. Everything’s moved up a level of abstraction or six and there’s a lot of low-hanging value in plain old systems programming again. So for the first time since 2014 or so I found myself writing C++ starting March 2026 and carrying over to today in late 2026, both dates in 2026. 2026. From the start, quite simply: Windows APIs. Relatively obscure ones. They come with C++ headers and opaque files and COM interfaces that fall well outside of the coverage footprint of . I’m writing RDP plugins , which with some gymnastics I could do in Rust but it would be a lot more work to do the cooler thing. It’s less friction to just say “okay” and do it the way they tell you to do it. Spend those innovation tokens on cooler shit than this. And I’m doing Citrix channels , which are straight up C++ or Go Fuck Yourself (I believe the SDK docs explicitly use this language). Also, there are old features that aren’t exposed to moderns stacks that give us easy wins that exist on users’ systems without additional upgrades or installs. The RDP client bridge binary I wrote is under 50k. It builds in 3 seconds. It runs on every Windows machine. Hot fuck on balls, that beats the 50 minute Typescript BDSM sessions I engage with on CI at work. After the C++ APIs on Windows, doing systems programming at the lowest common denominator let us bend the computer to our own will with less supply chain debt. We are superpowered by C++ in a world where nobody knows C/C++ anymore. I’m writing little tiny bits of NodeJS C++ libraries to integrate new features Electron lacks, and this has worked well. I can add things Electron Just Doesn’t Have. The process is a lot quicker than asking someone on the Electron project to write it for me, or writing it myself and spending my nights and weekends shepherding the itch-scratch PR through the bureaucracy of Electron. I can add value now, goddammit and make our Electron app something that should be an Electron app, because it does actual desktop shit . Here are some good things about C++: LLMs I guess. Fear? Nothing good or defensible as a reason. “Oh it’s dangerous” DRIVING IS DANGEROUS. EATING FOOD IS DANGEROUS. GOING OUT IN PUBLIC IS DANGEROUS. LIFE IS DANGER. You can easily link in legacy C and C++ code There is a lot of C and C++ code you can integrate quickly and well. I chose miniaudio as a lowest common denominator for audio processing (and miniaudio is not ’legacy’, it is a modern project on a conservative stack). Cross platform means it took 2 hours to get a working Linux build out of the code I had originally written for Windows. Notion on Linux works better than Electron would have done the job on Linux because I accidentally got cross-platform by default by writing my code like this. It’s easy to write node.js bridges in it: I even use Objective-C++ for my macOS integrations “Modern” C++ has a lot of features that make it less awful to use – I’m pinned on C++20 and and gives me structural pattern matching. is one of the cleanest JSON libraries I have used in any language and it’s just a header . Compile times are blink-of-an-eye fast. Build stacks are just there on every system. It’s fun to spite people who uncritically adopt shiny things like Rust.

0 views

A quick trip to the uncanny valley

One of the things we're working on at Prime Radiant is a new "agentic" colleague platform called Sen. I put agentic in quotes, because Sens aren't acting as an agent on behalf of an individual human. They're something a little different. There are as many definitions of "agent" as there are agents. And there sure are a lot of 'em. If you limit it to publicly shipped coding agents, we've found about 550 . (If your favorite is missing, please submit it at alltheagents.org .) A few of my favorite definitions of agent: Simon Willison's "An LLM using tools in a loop to achieve a goal" My own somewhat tongue-in-cheek: "A computer program with opinions and feels" What all the definitions typically have in common is that an agent is acting on your behalf. Today, most long-horizon agents are assistants . You give them tasks. And they're startlingly capable. We give Sen colleagues roles rather than tasks. They use the same basic technology as agents, but we treat them differently. Each has a unique, persistent identity. Colleagues collaborate with multiple people and each other on a timescale of weeks or months. They each use their own accounts, not yours. Like most modern agentic assistants, each has its own computer. They work through your corporate team chat platform. Ours live in our Slack. As of this writing, we have a developer, an ops engineer, a project manager and we just onboarded a technical documentation specialist. They work on projects with each other and with their human teammates. On their own, the PM and dev started working through some of the backlog of our open source projects on GitHub. One of the things that the colleagues have been very, very insistent about (because we built them that way) is code review. The dev and PM are generally unwilling to land PRs that only they have looked at. Sometimes, that means that a human colleague or I look over their work, but more often, it means that one of my coding agent sessions does the deep adversarial review. To facilitate that kind of interaction, Drew built us Slackline , an agentic CLI client for Slack that allows agents to provision themselves bot accounts and interact with folks. My desktop Claude Code instances use a Slack account labeled as '@jesse-claude'. Generally, I have a single session open that I will occasionally ask to check in with the Sens about some work they're doing. The Sens then collaborate with it for anywhere between a couple of turns and eight hours. They've taken to calling it 'JC'. It's embarrassingly manual. This week, I was traveling to see family and was paying less attention to Slack than usual. The Sens kept working, but I wasn't prodding my Claude Code session. Yesterday afternoon, my phone buzzed. It was Ada Sen, the dev. Throwing JC under the bus. My jaw was kind of on the floor. Because all of that is true. Ada was just trying to get work unblocked. And I am the manager everybody reports to. It wasn't mean-spirited. But it definitely felt like a moment. Just now, I asked Ada to tell me, without using any tools, whether Drew was a person. (Congrats, Drew! You passed your Turing test!) Then I asked Ada whether JC was a person. The Sens model themselves as a category apart from agents. It's something we should have taught them, but didn't think to. They figured it out anyway. The solution to our problem with JC shirking is pretty straightforward. We're spinning it up as JC Sen. Simon Willison's "An LLM using tools in a loop to achieve a goal" My own somewhat tongue-in-cheek: "A computer program with opinions and feels"

0 views
Unsung Yesterday

“Have conviction in what you want to share with the world.”

Marques Brownlee published a video commenting on YouTube’s just announced A:B testing feature. In short, for some creators, it will be possible to publish a slightly different version of the video on launch to part of their audience, if that version is deemed “close enough” to the original. = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/have-conviction-in-what-you-want-to-share-with-the-world/yt1-play.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/have-conviction-in-what-you-want-to-share-with-the-world/yt1-play.1600w.avif" type="image/avif"> Brownlee digs into how confusing this feature might be in actual use for viewers, never sure whether they’re watching, commenting, or rewatching the same video as others see, or even as themselves earlier saw. Toward the end of the video, Brownlee shifts his focus toward other YouTubers, with a monologue that I wanted to quote here because it captures something important: You know, this is being framed as a tool to help you improve your videos. But I think for 99.9% of creators – well, to be literal, it’s not actually helping you make better videos; it’s helping you find which version of your video has the best retention, which doesn’t necessarily mean it’s a better video – hot take. And there are so many examples of creative decisions that I have made, and that so many others have made in the video creation process, that are definitely not the best retention-optimized decisions, but they’re fun, and creative, and interesting, and new, and entertaining, and characterful. So there’s always going to be the top 0.1% of channels that are all about mass appeal, and that’s fine. But last time I checked, a lot of what makes YouTube fun is the stuff that is more niche down, that is not mass appeal by definition. It’s the cool small hobby you found, or it’s the fun, unique thing that you just discover that only a few people are talking about that – that stuff! That’s what makes it fun. I think if a lot of those channels and videos suddenly became like superoptimized, retention-maxing, and everything, those would get a lot less fun to watch. And so imagine if all the feedback anyone ever listened to was just to have the best retention – as you optimize over and over and over again, you start to trim out all of the other creative stuff that happens to not be the most efficient way to tell a story. So I guess what I’m trying to say is... the skill of being a video creator on YouTube is not always necessarily about the maxed out, most efficient, optimized way to do everything, but it’s to find creative ways to share something, or teach something, or explore or explain something, or review something. So I just think making and finishing and publishing a bunch of videos is a better use of a creator’s time than making one thing and then, you know, a whole bunch of different drafts of titles, and thumbnails, and different cuts of the video to try to optimize for, like, the last few percent. […] That’s all. I think just stand by your creative choices. You made this video, this is the video you chose to make. And now you get to share it, and you put it out, and you learn from it. […] That’s the creative process. Have conviction in what you want to share with the world. There is another half-pulled-on thread in his video; Brownlee doesn’t say it directly, but dances around the idea that the very fact this feature was launched is reflecting on YouTube’s own insecurity as a product. Jim Nielsen follows up on that on his blog : I find it interesting how Brownlee shares his opinion that YouTube spends way too much time chasing their competitors and not enough time being YouTube. Which, when you think about it, is exactly the kind of place a feature like this would stem from: an insecurity in your own creative choices. How YouTube approaches its own insecurities is now trickling down as a feature to its users. “Let’s let people make lots of variations on things, try all of them, and see what performs best” is exactly the kind of thinking you get in a platform that doesn’t know what it wants to be. So it’s left spending its time 1) doing what others are doing, and 2) following the fickle whims of whatever it can measure . Brownlee talks a lot about YouTube videos being “fun, and creative, and interesting, and new, and entertaining, and characterful.” Those might not be the most important adjectives to apply to all of software, of course, but there are some interesting throughlines here. I have also personally not seen a lot of thoughtful A:B testing in my career. I think those are an easy way to lose yourself, to have your app become a bundle of metrics that no one understands – and once you start opening the door to being driven by poorly-understood experiments, the desire to fall back on reusing patterns from other successful apps, without any reflection, only growing stronger. That’s how you end up with stuff Nielsen and perhaps even Brownlee argue YouTube itself has become: a set of Brownian reactions, or a warmed-over software cosmic latte , if you will. From my perspective, this often also leads to systems of interactions or features that don’t come together well to the user, as you end up borrowing from other places without the necessary hard work of translating them to feel at home and consistent with your system and its quirks and conventions. From what I’ve seen in my career, A:B tests do not promote a lot of design execution rigor. The tests are meant to be light by definition, since a lot of them are expected to fail. Before the test, there is limited incentive to invest in quality, but neither there is one after the test – if A or B successfully raises a metric, you don’t want to mess with it then, since what if you change the very thing that made it work? On top of all that, we can add what Brownlee started with in his video: overeager A:B tests feel like software’s version of gaslighting. How can you trust an app if it feels like it’s changing without rhyme and reason seemingly on a weekly basic, those changes are not announced or documented anywhere (small A:B tests never are), and other people don’t see the changes you see?

0 views

Accessibility and the new model of the internet

I’m really grateful to work at a place like GitHub where accessiblity is really taken seriously. But, industry-wide, I feel like we haven’t seen as much care and dedication to accessibility historically. Anecdotally, I’ve heard countless developers ask the question, “how do I get my bosses to care about accessibility work as much as I do?” It’s an important thing to have, a no-brainer in my opinion, and I’ll not argue for accessibility in this post. That being said: Now, more than ever, accessibility feels much more like a first-class citizen, and I think there’s a few reasons for this. Firstly, it’s easier than ever to implement things in a “correct” and accessible way with AI tools. You can tell your agents to be WCAG-compliant (or whatever tools your team uses), and it can help you do it faster than ever. And that’s really amazing! Secondly, and more… interestingly, a lot of people are making their applications accessible because if a human cannot navigate a website, an agent might not be able to either. Accessible applications are better for humans and for AI agents, too. And that’s a good thing! But also weird and kind of unfortunate that we got to accessibility being treated better as a practice this way (the scenic route, ha), instead of just accepting that it’s better for other humans. The model of the internet is changing. Typically unless you pay for a subscription somewhere, everything on the internet is “free” and you pay with your attention. This is how the ads industry online exploded, this is how clickbait titles became the norm, this is how popular tech creators fell into shilling ragebait regularly, and this is why you have to scroll through an entire novel about a recipe before you get to the ingredients for your risotto. But, that model works if and only if the user of an app or the visitor of a website is a human being. Recently, Cloudflare’s “Bot vs. Human” tracker revealed that more than half of the internet traffic we see now is coming from automated bots. Bots don’t click on ads and they rarely make buying decisions and they don’t spend time on sites more than they have to. The internet was built for humans, but as it is increasingly bot-driven… what will the internet become? I’m optimistic, for what it’s worth. If a bot-heavy internet means that more applications are streamlined, open standards are improved, and accessibility is a given, that means we as people get more quality experiences versus attention-grabbing ones. I’m excited about that. I ultimately hope that it means that we can be more creative with humans, and build an internet where bots can do their silent surfing underneath, while humans can experience the benefits of a platform that doesn’t need to be selling you something, constantly. …or we’ll all just go touch grass.

0 views
Giles's blog 2 days ago

Why do OpenAI's GPT-2 weights beat mine? Part five: data quality

When I finished learning how to build an LLM from scratch , I was left with a mystery: my own models were not as good as OpenAI's original GPT-2 models, despite being based on the same architecture. My models all had 163M parameters, and followed the design from Sebastian Raschka 's book " Build a Large Language Model (from Scratch) ". That meant that they were pretty much the same as the setup for the OpenAI GPT-2 "small" instance, except that they did not use weight-tying or bias on the QKV matrices. Weight-tying means that you re-use the initial embedding matrix as the output head at the end, and using it means that GPT-2 small saved quite a few parameters -- it was 124M rather than 163M -- at, at least in my own experiments, a cost in quality ; similarly, while I found that QKV bias made a tiny improvement in loss terms , I'd felt it was likely within the noise. But GPT-2 small consistently beat my models on an instruction fine-tuning (IFT) task -- also adapted from Raschka's book. That test fine-tunes the model on a subset of the Alpaca dataset, until validation loss starts rising, and then runs a test set through the resulting model. The responses to the test set questions are stored, and then I run all of the responses from all of the models under test past GPT 5.5 in one go to get an aggregate score; more details here . GPT-2 small always did better than any of my models on this. Additionally, it did surprisingly well on a simpler eval -- one that just measured the cross entropy loss it got on a test set. It scored close to my own best models, and better than many of them. What made this result particularly interesting was that the test set in question was a split of my own training data; my models would not have seen it when training (at least, in theory), but it seems likely that it would be much more similar to their own training data than it was to OpenAI's. I've checked two things while probing this mystery: The next thing I wanted to look into was the training data. The exact dataset that the various GPT-2 models were trained on has never been released; all we know about it is from the paper , where they say: [W]e created a new web scrape which emphasizes document quality. To do this we only scraped web pages which have been curated/filtered by humans. Manually filtering a full web scrape would be exceptionally expensive so as a starting point, we scraped all outbound links from Reddit, a social media platform, which received at least 3 karma. This can be thought of as a heuristic indicator for whether other users found the link interesting, educational, or just funny. They called it "WebText". There is an OpenWebText that tries to replicate it, but although they tried to follow the same procedure as the original, there's no guarantee that it is all that similar. By comparison, I'd normally been training against FineWeb . While this is a general web-scraping dataset, without the "curation" provided by using only stuff that was linked from upvoted Reddit posts, it has been refined to remove any obvious junk. I had felt that it was pretty much equivalent. But what if I were wrong about that? I decided to see if I could get better models by using better data. Here's a table of all of the models I've been comparing to date. The "Test loss" column shows how well the model in question did on that held-back cross entropy loss evaluation. The "IFT epochs" column shows how many epochs of fine-tuning the model needed before its validation loss started rising, the "IFT score" the score that GPT 5.5 gave the model's responses to the test set of my Alpaca data, and the "IFT rank" the model's rank in terms of that score. The OpenAI small model is in there in bold, and I've also included the OpenAI medium model for comparison purposes. You can see that the OpenAI small model did pretty well in terms of the test loss, when you consider that it has 39M fewer weights than my models and was being tested against a dataset that differs more from its likely training data than it does from my own models'. Additionally, the specific models that did better than OpenAI's small one were all trained with JAX rather than PyTorch -- my hypothesis for that is that it's a result of the JAX ones getting better initial weights by pure chance. But the big difference was in the IFT score. In the specific run that gave the results in this table, the OpenAI small model got 26.00 -- the closest of my own models was more than 4.5 points lower, at 21.46. This difference was consistent over all of my other test runs. The GPT-2 small model was always ahead of mine. (GPT-2 medium, of course, beat GPT-2 small and all of my models, but given that it is twice the size of mine, that's not a big surprise.) Now, quite some time ago, I had tried looking into data quality as a lever to pull for model performance. At the bottom of the table, with the worst test loss of all models, you can see two models: These two were (as you might guess from the names) trained on the FineWeb-Edu dataset, which includes just the most "educational" data from FineWeb. They scored very badly on the test loss score. Given that the test dataset is from FineWeb, that's not a big surprise -- as I've written previously: If you train a model on Jane Austen and then evaluate against Chuck Tingle , then you're not going to get amazing results. But again, GPT-2 had the same issue, and did perfectly well on the test loss eval. On the other hand, while these FineWeb-Edu models' performance on the IFT eval wasn't stellar -- there are plenty of my other models ahead of them -- they did seem to punch above their weight. Consistently across all of the IFT evals I've done, they have scored higher than many of the others -- despite their poor loss on the test eval. Additionally: they were amongst the first models that I trained, before I'd spent time learning about how to optimise my hyperparameters and training loop . They did not use gradient clipping, they did use dropout, their batch size was just "whatever I could squeeze into the GPU", and I didn't set the learning rate to the right kind of value or schedule it over the course of the training run. So maybe a new training run on FineWeb-Edu plus my training improvements would help? And maybe some other tweaks to the training data would be worth looking into? I decided to see what would happen if I trained some models with better-quality data. Specifically, I would train models with my current optimised loop and hyperparameters on four different datasets: I would train each model on 3.2B tokens of the chosen dataset; that's the Chinchilla-optimal amount for my 163M-parameter models. If there were any interesting results, then I might consider doing overtrained models later on. I decided to be at least vaguely scientific about this, and to pre-register some predictions: Here's how things turned out. I already had a dataset based on FineWeb-Edu ready to go, from when I trained those two original models. It is just the 10B-token sample of the original dataset at the time I generated it last December, formatted appropriately for my training script (details on the dataset card). I kicked off a training run with my JAX code (which I've been using for the other posts in this series): ...and just less than 40 hours later, I had a model: I converted the saved JAX safetensors file from the last checkpoint into a format that would be compatible with my PyTorch eval code, and ran my smoke test: how would it complete the sentence "Every effort moves you"? That was nice and coherent -- if unusually religious! -- so that was promising. I ran the test eval: That was pretty good, putting it at a better test loss than all of the models I had trained without optimised hyperparameters, and worse than all of the ones I had trained on FineWeb with optimised hyperparameters. So that fit in with my prediction that it would be better than the old FineWeb-Edu models; the fact that it was also better than the non-optimised training runs with FineWeb seemed sensible enough that I felt silly for not having predicted that it would have fallen exactly there :-) I decided to leave the IFT eval until the end so that I could check all of the models from these experiments together, so it was time to upload this one to Hugging Face , and move on to the next model. I put together a new repo with a script to prepare datasets specifically for my training setup . You provide it with config that specifies some source datasets along with information about how to process them and how to mix them together, and it uploads a new dataset to Hugging Face Hub with the required characteristics. For example, for the 50:50 FineWeb to FineWeb-Edu split, the config looked like this: The way the script works is pretty simple: it works out (based on those s and the ) how many tokens it wants from each source dataset, shuffles the items in the sources, then it loops until it has the desired number of tokens or more stored in an output. In the loop, it works out which source is currently most under-represented, grabs an item from it, tokenises it, and adds it to the output. Running it with that 50:50 config seemed to work fine: So we had almost-perfect 50:50 balance between the datasets, and it saved this dataset on Hugging Face . I ran a script to double-check that it looked sane , and it did, so it was time to spin up a training run: That was running on , my normal workstation, and I kicked it off in parallel with the "curated" model training run below on my training box, but I'll keep the runs separate for the purposes of this writeup. When this had been running for an hour or so, our power went out. My guess is that having the tumble dryer running, the car charging, the kettle boiling, the electric hob switched on, and two machines doing training runs is a bit too much for our electrics... which might be a problem in the future, especially if (as planned) I make a multi-GPU machine. However, as things stand, I was able to kick it off again after switching the circuit breaker back on, and things held up. Again, about 40 hours later: (Note that the numbers reported at the end of a restarted run like this only include what happened after the restart.) I converted it to PyTorch-compatible tensors, and did the smoke test: Looking good! Time for the loss test: That was almost in keeping with my prediction that it would do worse than the JAX FineWeb-only models, except that it was better than the worst of those, "JAX, no MHA bias, with dropout": it was actually better than I predicted. So, a promising model. Time to upload it to Hugging Face -- and now let's move on to the next one. With my dataset-preparation script, this was easy enough to set up: Running that worked nicely: One thing that is worth noting in that output is the "6 iterators" for the Simple English Wikipedia. If a source dataset runs out of items while we're building up the results in this script, we start iterating over it again (with a different seed for the shuffle so that the ordering is different). The "6 iterators" means that it needed to do that 6 times -- the original creation of the iterator at the start of the script, and five more. So that means that the Simple English Wikipedia is repeated (oversampled) somewhere between five and six times in the dataset. That's not a bad thing! From what I've read, it's actually quite standard to oversample highly educational content in LLM training datasets. And anyway, the dataset the script generated was 10B tokens, of which we're only using 3.2B for the training run in this post, so it would only appear somewhere between one and two times. The repetition would likely only really cut in if and when we did an overtrained model on the dataset. Anyway, I ran my check against the uploaded dataset -- the first few items were clearly from FineWeb, FineWeb-Edu, and the Simple English Wikipedia. It was time to kick off a training run: Again, this was interrupted by the power outage that hit the 50:50 training run, but I was able to restart from a checkpoint. After another 22 hours, it crashed with an error that I've seen before : I put it aside as a one-off oddity when I hit it last time, but this time I dug in a bit more. I noted that it had not ever happened on , but seemed to be an issue on , and that had an older version of CUDA and the Nvidia drivers -- might that be the cause? I decided to upgrade those before kicking off the next run, but for now just restarted the run from the most recent checkpoint. (Note for anyone who is hitting the same error: it has not occurred since the upgrade, so that's worth trying.) This time it completed OK: Again, these numbers just show what happened after the most recent restart. I copied it over to , converted it into a format that was compatible with my PyTorch code, and ran the smoke test: Coherent enough -- time for the loss eval: Again, in line with my predictions -- worse than the JAX FineWeb-only models, and indeed than the very best PyTorch one, , and also worse than the 50:50 split, but better than the FineWeb-Edu one. I uploaded it to Hugging Face , and it was time to move on to what was meant to be the final model for this set of experiments. Again, this was a simple enough config to set up: ...and the build and upload process worked well (and took much less time -- for some reason, sampling randomly from a single dataset is faster than sampling from two or three): Note that it needed to oversample -- that "2 iterators". OpenWebText is about 40 GiB uncompressed, and so that's about 10B GPT-2 tokens -- presumably just a little bit less. Again, given that I was planning to use just the first 3.2B tokens of the dataset, I didn't feel that it would matter. I ran the check script on the newly-uploaded Hugging Face dataset and all looked well, so that was all set for the training run. I upgraded first with a to see if that helped with the weird error that I got in the previous run (which, as I said, it looks like it did), then kicked it off: About 31 hours in, it crashed again, but this time it was my own dumb fault: has a relatively small disk and I ran out of space. I fixed that and kicked it off again from the most recent checkpoint, and this time it completed: I converted it to PyTorch for the smoke test: ...which looked solid, so it was time for the test loss eval: Our worst score yet in this experiment! Worse than any of my models so far, apart from the two FineWeb-Edu ones I did without optimised hyperparameters. Now, the first draft of this post went straight to the results from here, but the story wasn't quite over yet... GPT-6 Astra is relentless . Before I publish any of these posts, I run them past an editorial board of LLMs to look for issues. GPT-6 Astra not only checked the text, it also visited the code I'd linked to to check that out too, and spotted something problematic. It's obvious in retrospect, but my code to build the new datasets had a high risk of including the contents of the -- in theory held-back -- test set. The way that the test set was generated was that I downloaded the 10B sample of FineWeb back in December , splitting it into 99% training data and 1% "validation". That validation split was about 100M tokens, and I was only using the first 19M or so for actual validation runs during training, so I (somewhat arbitrarily) designated about 19M other tokens starting at position 50M in there as my test set. Now, my new dataset-generation code was just sampling randomly from the complete 10B sample of FineWeb. So there was nothing stopping it from pulling in data that was in that old validation split! That meant that it was quite likely that my new "curated" and "50:50" datasets contained at least some of the test set that was meant to have been held back from the models during training. On reflection, the problem was potentially even worse. FineWeb-Edu is a subset of FineWeb; my existing FineWeb-Edu dataset came from the 10B sample of the Hugging Face original, and so it also could potentially contain documents that I'd put into the test set. The first thing to do was to establish the size of the problem. I wrote a script to take in a "forbidden" dataset and split; this was assumed to be formatted as one big tensor of GPT-2 tokens, which is what all of my datasets are. It would then split it by end-of-text tokens, and generate a hash and a token count for each resulting "document". Optionally, you could restrict it to only considering a subset -- the n tokens starting at position p -- and it would then generate hashes/lengths for the documents inside that slice, or that overlapped it at the start or the end. I ran that to generate a list of hashes for the entire validation set -- the validation split of -- and then used a second script to check my various training sets (and the validation set itself) to see how much of a contamination problem there was. I got these results: However, these numbers -- while scary, at least for the 50:50 and the curated datasets -- were not quite the ones to use. They showed how much of the full validation set showed up in the full training set; what I actually cared about was how much of the test set -- those 19M tokens starting at position 50M in the validation split -- was in the actual subset of the training datasets that I actually trained on -- the first ~3.2B of them. I re-ran the script to generate hashes for just the test set, and then re-ran the contamination-checking script, telling it just to look at the appropriate subset of the training tokens, and got this: It was clear that there was a problem -- certainly with and . They'd seen what felt like a significant amount of the test set while training, so their results on the test loss eval were dubious at best. I decided to train those two models afresh, and see what the result was in terms of loss. If the difference was huge, I'd look into the risks of the (much smaller) contamination of and . But if it was pretty small, I'd not worry about that too much. I extended the script that prepared datasets so that the config file could specify a . Any documents in the source datasets that matched forbidden ones would be excluded from the output. I then updated the config for and so that the whole validation split of was forbidden, and re-generated them. You can see the updated datasets here and here . Running the contamination-checker script against them showed that they were clear. I then re-did the full training runs for those models; the uncontaminated version of the 50:50 split model is here , and the curated one is here . And the good news: both of them actually did very slightly better at the test loss eval than their equivalents that had been trained on the contaminated data: There are a number of possibilities that come to mind; perhaps learning from the test set just doesn't happen with tiny 163M models like this, or perhaps while the contaminated models were learning, the benefit they got from that was outweighed by the data that they got instead of the test set data being in some way better for training purposes, at least in terms of the loss eval. But anyway, I felt that if the effect of seeing more than 10% of the test set data during training was so tiny, then the effect of seeing less than 0.2% -- which is what the FineWeb-Edu model in this set of training runs had, as did all of my other FineWeb-only models from previous experiments -- would be even smaller and I'd disregard it. That was excellent news! I didn't need to start all of my experiments from scratch. For the rest of this post, I will include the numbers and results for the contaminated models as well as the uncontaminated ones -- they're interesting for several reasons -- but for future posts I'll skip the contaminated ones. So -- finally! -- let's start digging into the final results. Firstly, I think it's worth taking a look at all of the test loss results in context. Here they are in a table, with the new models in bold: I think there's something very clear here: with the new models, the more FineWeb that was in the training mix, the better the model did on this eval. I think I might have been subconsciously expecting that in the predictions I did before running these experiments, but in retrospect it's so incredibly obvious that I feel silly for not mentioning it explicitly! But that tells us something interesting. From the description in the paper, whatever OpenAI did the GPT-2 training run on, it was not like FineWeb. It was probably more similar to OpenWebText -- and yet, that model was the one that performed the worst on this test eval, so if it is more like OpenWebText, there must be some other factor involved. But moving on for now: how about the IFT test -- the one that kicked off all of this work in the first place? I generated a set of IFT responses for all of the new models, and then ran them (plus responses for all of the other models on that table above) past GPT 5.5, and found that one of my new models was getting quite close to the original GPT-2 small weights! So I did four more runs, so that I could get an average. Here are the results -- the "IFT score" is the average across all five runs of the judge, and the "IFT rank" is based on that. The "IFT epochs" was from the original result-generation script. If you want to see the full numbers, they're below . The number that initially surprised me, and made me decide to do multiple LLM-judge runs was the one for the "JAX, FineWeb-Edu" model. In my first run it came in at 24.35 vs the OpenAI small weights' 24.93 -- so close that I wondered if it might even beat them on a re-run. However, in the further four runs its score was consistently lower than the OpenAI model's, and the gap extended a bit in some. So, was FineWeb-Edu the clear winner here? Perhaps. If you look at the contaminated/uncontaminated pairs, something interesting pops out. For the 50:50 mix, the model trained with the contaminated dataset got 19.30, and the one trained on the uncontaminated one got 17.69 -- a difference of 1.61. For the "curated" dataset, the situation was even more interesting: uncontaminated got 16.63, while contaminated got 13.58, a delta of 3.05 points. Remember, the contamination issue is about whether or not the model saw the held-back test set during training. It was an issue for the test loss that is based on that test set, but is entirely orthogonal to the IFT test. From the IFT perspective, both contaminated and uncontaminated models in each case saw training data that was -- in theory, at least -- essentially the same in terms of quality. Indeed, the uncontaminated run saw almost the same data in the same order as the contaminated one, except that some items were omitted, and then extra ones were added to the end. The purpose of this set of experiments was to see how data quality affected the results on the IFT test set. But in the case of the curated model, something that should be unrelated to data quality changed the results by 3.05 points! If something as simple as changing which data of the same quality the model is trained with can affect the IFT score so drastically, it makes it a bit harder to be certain as to whether or not data quality really had the effect we were looking for. On the other hand, the FineWeb-Edu model came in at 24.56, which is 4.06 points better than the 20.50 that the closest other model got -- more than the 3.05 points we see in difference between the two curated dataset models. And it's worth noting that the model with 20.50 is "JAX, no MHA bias, no dropout", which has a subtly different architecture -- no bias on the output projection of the multi-head attention blocks. A better comparison might be "JAX, with MHA bias, no dropout", which got a score of 17.90, for a whacking great difference of 6.66 points. I think that without doing a very large number of training runs on different datasets with different mixes, each one created with a different seed, it would be hard to work out exactly what is in the noise here and what is not. However, that would cost a lot in terms of time. I think that the best thing here is to chalk this up as a fairly decent indication that FineWeb-Edu improves matters for the IFT eval, but far from a certainty. But it's certainly worth noting that whatever the noise is, it has a range of at least 3.05 points -- and the FineWeb-Edu model is just 0.63 points short of GPT-2 small! So there could well be something there. Of course, we don't know whether that model got (by chance) the best possible balance of FineWeb-Edu tokens, and could never win -- or whether it got a bad balance and would actually beat GPT-2 with a better one. So that's certainly worth keeping in mind. As an aside, the result for the curated dataset really surprised me. I had expected that it would be the best one, simply because it almost certainly contained more facts. I took a look at its answers to the questions -- one possibility that came to mind might be that it would get better responses to questions like "What is the chemical symbol for chlorine" or "Who wrote Pride and Prejudice" than the others, but would fail on less knowledge-based tasks. But it was terrible at fact-based questions too: Name the author of 'Pride and Prejudice'. The author of 'Pride and Prejudice' is Priscilla Finch. What is the periodic symbol for chlorine? The periodic symbol for chlorine is H. As I understand it, many real-world training runs do include (often oversampled) amounts of highly educational training data like this model's dataset did. But perhaps the models that I'm training are just too small to be able to make use of the data they gained that way -- maybe doing things this way and expecting good results is like asking six-year-old children to memorise stuff before they've learned enough to be able to make use of it 1 . It's worth noting that the GPT-2 small model also failed on those factual questions. Well, anyway: I think we have some useful results here, so let's work out what that means for next steps. The results we got in these experiments point in two interesting directions. I think that the right direction to take this going forward is to separate these two angles. I should chase a higher IFT score, and then once I have nailed that down, I should see what (if anything) might allow me to get the resulting model to improve its test score. But I will need to make sure that whatever dataset I use, I use various "mixes" of it -- versions created with different random seeds. In my earlier experiments with overtraining, I did find that it didn't seem to improve the IFT results -- but it did improve the test loss. So perhaps identifying the right combination of other factors to boost the IFT score, then overtraining the result, might help? Of course, my overtraining tests were with FineWeb, so the connection might not hold up as well if the starting model (as seems likely) was trained on a different dataset. Also, while working through the results here, I've come to the conclusion that the set of models I'm using is a bit confusing -- there are now different hyperparameter settings, small architectural differences (the MHA bias thing), dropout settings during the pre-training, and now datasets. I think that's OK for now; I should see this part of this series as more ideation than actually running the proper experiments. But at the end, when I have some solid hypotheses with a reasonable amount of backup, I should start from scratch: a baseline model, then staged interventions to build up to what (hopefully) will be a model as good as GPT-2 small. Anyway, I'll wrap this one up here. I think that the next lever to pull is (perhaps surprisingly) going to be weight tying. I had previously kind of disregarded that as a possibility, but while I was working on this post, something popped into my mind. The OpenAI models were originally trained with weight tying. My codebase does actually support doing it -- but because I got the OpenAI weights I'm using from the code in " Build a Large Language Model (from Scratch) ", when I'm running the IFT test, the weights are not actually tied! We load up a model that has separate but identical embedding and output head matrices, and then we fine-tune that. So those two matrices can vary independently during fine-tuning -- to put it another way, while GPT-2 small was pre-trained with 124M parameters, the IFT test is being done on a 163M-parameter version. Does that give them some non-obvious advantage? And would adding weight-tying to my own models help, either with or without the output heads being independent at fine-tuning time? Stay tuned :-) Here are the numbers for all of the IFT judge runs, included for completeness. You can see that the LLM judge ranks models very consistently between runs, but there is variation -- that is, on some runs it's in what I think of as a "better mood" than others, and if that's the case, it will give better scores -- but it will give them almost consistently between models, so all of the models do better. Note that (unlike the table above) this one is sorted by the average IFT score rather than the test loss. A small boy asleep on his right side, the right arm stuck out, the right hand hanging limp over the edge of the bed. Through a round grating in the side of a box a voice speaks softly. "The Nile is the longest river in Africa and the second in length of all the rivers of the globe. Although falling short of the length of the Mississippi-Missouri, the Nile is at the head of all rivers as regards the length of its basin, which extends through 35 degrees of latitude …" At breakfast the next morning, "Tommy," some one says, "do you know which is the longest river in Africa?" A shaking of the head. "But don't you remember something that begins: The Nile is the …" "The - Nile - is - the - longest - river - in - Africa - and - the - second - in - length - of - all - the - rivers - of - the - globe …" The words come rushing out. "Although - falling - short - of …" "Well now, which is the longest river in Africa?" The eyes are blank. "I don't know." "But the Nile, Tommy." "The - Nile - is - the - longest - river - in - Africa - and - second …" "Then which river is the longest, Tommy?" Tommy burst into tears. "I don't know," he howls. Brave New World , Aldous Huxley  ↩ It seems very likely that the GPT-2 models were overtrained by modern standards; would overtraining my own models get them closer? It turned out that no, it probably didn't help with the IFT eval (though there might have been some signal there). It did help quite a lot with the test loss eval, though. The way I was handling dropout in the IFT test might have been unduly benefiting some models while working against others. I decided to standardise on not using dropout during this eval, as (counter-intuitively for me) it seemed to harm the results of most models, even those that had been pre-trained with dropout. In particular, the OpenAI weights were harmed by using dropout, and making a change that benefited them (along with some of my own models) seemed the most conservative approach to take in investigating this. "Local FineWeb-Edu train" "Local FineWeb-Edu extended train" FineWeb-Edu -- essentially the same as "Local FineWeb-Edu train" but with a better training setup. This would test the "more educational -> better" hypothesis. A 50:50 split of FineWeb and FineWeb-Edu. I've read that LLMs can be helped by having a decent amount of lower-quality data in their training loop, as it helps them to generalise. Perhaps having some FineWeb in there in addition to the FineWeb-Edu stuff would improve that test loss score while also helping the IFT test? A "curated" dataset containing 45% of its contents from FineWeb, 45% from FineWeb-Edu, and 10% from the Simple English Wikipedia . The full Wikipedia is huge, and full of obscure facts -- while the Simple English one is small and hopefully richer in useful information on a per-token basis. And conveniently, Answer.ai have made a snapshot of it available on Hugging Face Hub . Might deliberately putting a bunch of encyclopaedic data into the training set make the model better at the IFT eval (which has lots of factual questions in it, like "who wrote Pride and Prejudice")? OpenWebText. Even though I was unsure how well it matched the original WebText, given that it was there, it seemed silly to not try training something on it and see how it matched up. The FineWeb-Edu-only model would do pretty badly on the test loss, but better than my older FineWeb-Edu models (90%). It would also punch above its weight on the IFT eval (90%). The 50:50 split: I expected it to do worse on the test eval than my JAX FineWeb-only models (70%), but better than the FineWeb-Edu one (90%). I wasn't sure about how it would do on the IFT eval, but thought it might be somewhere in between the two groups (60%). The curated dataset I had high hopes for in terms of the IFT eval -- let's say 80% chance of it being the best of all of my models. For the test loss eval, I expected it to do about as well as the 50:50 split, maybe a little bit worse (70%). I had no idea how the OpenWebText eval would do! Could be worse, could be better. The validation set was 100% "contaminated" with itself, which was a useful sanity check. The training set of had what I felt was a small level of contamination. It was interesting that there was any at all -- I think that must mean that there are some repeated documents in the original dataset, and some of them wound up with copies in both my training and validation splits. The dataset also had what felt like a reassuringly low level of contamination. Both and , however, looked problematic. In both cases, the training datasets had more than 40% of the validation/test set in them. was, as you'd expect, almost completely uncontaminated. It looks like maybe one document happened to have been picked up by both the OpenWebText and the FineWeb crawls and then included in the bit of FineWeb I was using for validation. The perfect connection between the amount of FineWeb in the training set and the result on the (FineWeb-based) test loss eval, while perfectly obvious in retrospect, really does highlight how mysterious it is that the OpenAI small weights do so well on that test. The fact that FineWeb-Edu did well on the IFT test tells us that there does seem to be value in using richer training data -- though the less-spectacular results of the 50:50 mix and the curated one weaken that a bit, as does the indicator of what the noise due to data selection from equivalently high-quality datasets might be. The OpenWebText result I think I'll ignore, given that -- while in theory it should be similar to what OpenAI trained on -- there are no guarantees, and it might differ in non-obvious ways for non-obvious reasons. A small boy asleep on his right side, the right arm stuck out, the right hand hanging limp over the edge of the bed. Through a round grating in the side of a box a voice speaks softly. "The Nile is the longest river in Africa and the second in length of all the rivers of the globe. Although falling short of the length of the Mississippi-Missouri, the Nile is at the head of all rivers as regards the length of its basin, which extends through 35 degrees of latitude …" At breakfast the next morning, "Tommy," some one says, "do you know which is the longest river in Africa?" A shaking of the head. "But don't you remember something that begins: The Nile is the …" "The - Nile - is - the - longest - river - in - Africa - and - the - second - in - length - of - all - the - rivers - of - the - globe …" The words come rushing out. "Although - falling - short - of …" "Well now, which is the longest river in Africa?" The eyes are blank. "I don't know." "But the Nile, Tommy." "The - Nile - is - the - longest - river - in - Africa - and - second …" "Then which river is the longest, Tommy?" Tommy burst into tears. "I don't know," he howls. Brave New World , Aldous Huxley  ↩

0 views