Latest Posts (20 found)

📝 2026-08-01 15:28: Our new addition to the family, Nelly, has arrived and she's causing havoc already. 🤣

Our new addition to the family, Nelly, has arrived and she's causing havoc already. 🤣 Thanks for reading this post via RSS. RSS is ace, and so are you. ❤️ You can reply to this post by email , or leave a comment .

0 views

📝 2026-08-01 13:46: Signed up for a Vinted account. My options to "secure my account" were adding a...

Signed up for a Vinted account. My options to "secure my account" were adding a phone number, or to logout. I fucking hate this side of the internet. Scumbags. Thanks for reading this post via RSS. RSS is ace, and so are you. ❤️ You can reply to this post by email , or leave a comment .

0 views
Taranis Today

Amnesty International's Taken-Down Article on the UK Anti-Rights Movement

Retrieved from: https://archive.org/details/report-a-growing-threat-the-anti-rights-movement-in-the-uk-july Since it's quite likely that archive.org will be the next target, I strongly suggest that people download and republish this file as widely as possible. For what it's worth, I believe Amnesty International wrote the original in good faith, and published it because it was the right thing to do. That they caved to pressure from an onslaught of lawsuits funded by the deep pockets of a billionaire (who has a personal beef with Amnesty for other reasons, but I'll not go into that here) is not to Amnesty's credit. I believe that their retraction was made under duress. Let the Streissand effect roll.

0 views

📝 2026-08-01 08:47: Today is the first of August. Where the hell did the first two thirds of...

Today is the first of August. Where the hell did the first two thirds of 2026 go??? 😵‍💫 Thanks for reading this post via RSS. RSS is ace, and so are you. ❤️ You can reply to this post by email , or leave a comment .

0 views
Unsung Today

You only get one chance to make a zeroth impression

Contra the previous post , just spotted this “Same settings, new look!” callout in Firefox, on a page I have never visited before, talking about a feature I have never used: = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/you-only-get-one-chance-to-make-a-zeroth-impression/1.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/you-only-get-one-chance-to-make-a-zeroth-impression/1.1600w.avif" type="image/avif"> I mentioned this once – when you get someplace new, everything is new. This redesign announcement’s conditions were engineered poorly, because they didn’t even consider people who didn’t use the feature before the redesign. I keep thinking of the very old computers from the 1940s, 1950s, and even 1960s, where the challenge wasn’t what to do with the computer, but with keeping the computer running. There was a machine called JOHNNIAC which gained a nickname “pneumoniac” because it refused to work on a warm day. On many other computers, there was a daily ritual of finding which of the vacuum tubes blew up and had to be replaced. Some teams had a literal downtime tax; they knew and communicated in advance that 40% of the days the machine would not be up to the task, consumed by tweaks and repairs to bring it just to the state of being ready to do work . When I look at those kinds of pop-ups and callouts, I think how we’re unwillingly recreating the very same conditions. I deeply believe people are exhausted not just by changes in software , but also constant questions, pop-ups, unnecessary notifications, “STOP 2 END,” emails to unsubscribe from, unsubscribes that don’t work the first time, and so on. Software should not require the user to occasionally click on buttons just to keep it running. #attention #onboarding

0 views

Weighing the Business Case for Quality

Most of us want to do great work. Building things of exceptional quality – work that’s beautiful, lovable, fast – is one of life’s great joys. However, the story goes, capitalism doesn’t value quality. Businesses don’t fund software that is lovable. Thus, you need to make a career choice: do you want to be a craftsperson, or do you want to make money? This dichotomy is nonsense. I mean, yes, these two things are in some tension. You can easily make money maintaining EOL plumbing if you can motivate yourself to do it well, and you can easily make beautiful products if you can afford to give them away for peanuts. But there absolutely exist businesses that are very large and successful in part because they invest in quality. Beauty, speed, and user experience can drive profits. Patrick Collison spoke about this back in 2023: My intuition is that more of Stripe’s success than one would think is downstream of the fact that people like beautiful things, and for rational reasons. Because what does a beautiful thing tell you? Well, it tells you the person who made it really cared. And you can observe some superficial details, but probably they didn’t only care about those and did everything else in a very slapdash way. … If someone really convinced me that the financial returns to craftsmanship and beauty – which I believe are greater – are not greater, I’d still do it. Life’s too short. So you could try working at Stripe. Even within Stripe, however, there are products and areas that don’t justify the level of care Patrick describes here. To get a real mandate for quality – to do great work, and earn a good income doing it – you need to ask yourself a question: is there a business case for quality here ? If you want the resources necessary to build something well, then choose to work on things worth doing well. There are a lot of reasons a product could be worth doing exceptionally well. Here are some: There are many others, but even just the above attributes describe some of the most craft-centric companies out there: Linear, Figma, Stripe, Wealthsimple, Apple, and so on. The economic details vary, but if a large company is spending the surprisingly large amount of resources it takes to build exceptional products, there’s a business reason to do so. Meanwhile, many of the most slapdash products out there are the inverse of this dynamic: imagine an internal control panel that an enterprise contract obliges you to deliver. A project like that shouldn’t be laboured over – it should be built to a minimum reasonable standard, then shipped. Its primary success factor is to exist. So if your goal is to work on things that are worth doing well, how do you figure that out a priori ? You need to get a sense of the business. Who are the customers, what makes them buy, and what makes them stick around? If you want to build a company that can justify making great stuff, ensure that you’re in a market where quality can confer an advantage. The analysis is a bit easier if you’re being recruited, since you can simply ask. Back when I ran Steamclock , occasionally I’d be approached by a prospective client that wanted to make a delightful app (excellent, that’s what Steamclock is all about) but for a use case that seemed… peripheral. Maybe it was an occasional-use product, or perhaps it was an internal tool. Sometimes it was plainly obvious that hiring an A-Team for the project was overkill, but other times, I’d explicitly ask: Thinking about what success looks like, how much do UX and quality for this product contribute to the success of the business overall? Sometimes they’d quickly realize it was a bad fit. But other times, they’d make a compelling economic case for craft and UX – starting us off on the right foot to execute on that vision. At the end of the day, there are countless ways to make a career building great software. The important thing to remember is that it’s not enough to care deeply about craft. It’s a great start, but if your customers won’t actually reward quality, you’ll be running uphill. Instead, work on products that have a business case for craft. You’ll have a durable mandate to do great work – and make money doing it. The product is used often – by many people, many times a week The motion is product-led – people self-serve, and refer by word of mouth The focus is retention – customers will vote with their feet The customers have high LTV – each buyer could grow into a huge account The buyers are users – quality impacts the decision-makers themselves The users are demanding – developers, designers, and other prosumers

0 views

Stateless MCP has recaptured my interest (and inspired mcp-explorer and datasette-mcp)

Tuesday was Stateless MCP day - the rollout of MCP 2.0, or the 2026-07-28 Model Context Protocol specification to use the more formal but less memorable name. This is the most significant change to the MCP spec since it first launched, and has also served to reignite my personal interest in the protocol. For background: MCP is the Model Context Protocol, which describes a standard way to expose new tools to LLM-powered agent frameworks. It was introduced by Anthropic back in November 2024 , had a huge spike of interest through much of 2025, and then became somewhat eclipsed by Skills (another Anthropic invention) when it became apparent that an agent harness with access to a terminal and could do most of what MCP did in a more flexible way. I wrote about that in my review of 2025 . I'm coming back around to MCP now. Giving an agent a shell environment with the ability to access the internet is fraught with risk , and requires a strong model that is capable of effectively driving such an environment. MCP tools are easier to audit and control, and simple enough that smaller models that run on a laptop can still drive them reasonably well. The new stateless MCP specification also greatly decreases the complexity of implementing both clients and servers for the protocol. I built three of those this week! The best demonstration of the difference between stateful and stateless MCP is in this May 21st blog post that introduced the RC for the new specification. It included a clear before-and-after example. The older stateful MCP (I'm going to call it "legacy MCP") required two HTTP requests - the first to initialize a session and obtain a , and the second to actually call the tool: The new stateless way uses a single HTTP request which looks like this: This is so much cleaner from both a client- and server-side implementation perspective. It's also a better fit for building scalable web applications, since now you don't need to maintain server-side state to keep track of those session IDs, or worry about routing the same session to the same backend machine. I couldn't find a great CLI tool for interactively probing an MCP server, so I had Codex help build my own. mcp-explorer is the result. It's a stateless Python CLI tool, so you don't even need to install it to try it out - it works with uvx like this: This queries Ade Oshineye's agentic-mermaid.dev demo MCP. The above command returns the following list of tools: Then to inspect a tool: This outputs a whole bunch of information, including the JSON schema of the inputs and outputs. To call that tool and pass arguments to it: Which returns: To get just the raw SVG try adding to that command. I got back this image : There are a few more commands in the README, but you get the general idea. I find building CLI tools like this to be a really productive way to get familiar with a specification, even if an agent writes most of the actual code. The second project is datasette-mcp , a Datasette plugin which adds a endpoint to any Datasette instance. This is probably the fourth time I've tried building this plugin, but thanks to the new stateless MCP specification I finally have a version that feels good to release. It provides just three tools: , , and . They do exactly what you would expect them to do - though is read-only for the moment. Wire these into an agent, or a chat tool like ChatGPT or Claude, and they'll gain the ability to run SQL queries against your hosted Datasette instance. So far I'm running it on the Datasette mirror of my blog, at datasette.simonwillison.net/-/mcp . It took a bit of fiddling to figure out how to attach that to ChatGPT and Claude, but I got there in the end. Here's a new TIL showing exactly how to do that. Here's a shared Claude session where I asked it: It ran 7 separate SQL queries to figure out the answer. My LLM tool is long overdue for an official MCP integration. The new alpha llm-mcp-client plugin is my attempt at exactly that: Here's the output (including reasoning trace, I'm using LLM 0.32rc2 ): Considering note count I see the question "count the notes" is probably asking me to tally up blog notes. It could also mean published notes or drafts, so there's some ambiguity there. I'll need to figure out the total number of notes, likely by querying the count for both published notes and drafts to get a clear answer. Let's execute that count! There are 151 notes . And the output of llm logs for that prompt. Once this is fully baked, I'm considering bringing it directly into LLM core. I'm excited to experiment with MCP in Datasette Agent and llm-coding-agent as well. A few months after MCP was first released, I wrote Model Context Protocol has prompt injection security problems , where I noted that the pattern of having end users mix and match tools pushed responsibility for avoiding data exfiltration attacks out to the users themselves. I hadn't coined the Lethal Trifecta yet, but that was absolutely what I had in mind. Then general agents with arbitrary shell and access came along, and that's so much harder to keep secure! Something I've come to appreciate about MCP is that it's much easier to reason about agent capabilities and what might go wrong than with arbitrary command execution in an open network environment - the default for most of today's general and coding agent tools. I plan to lean into MCP a whole lot more when I'm building sensitive applications on top of LLMs. You are only seeing the long-form articles from my blog. Subscribe to /atom/everything/ to get all of my posts, or take a look at my other subscription options .

0 views

Linkception

So many links in one post . I ended up going down all kinds of rabbit holes off the back of this single post (also the second time I've linked to Sal's blog today 🙃). I discovered Coyote's blog , and Sylvia's . So went ahead and read some of their posts. I was already aware of Brennan's fantastic blog , but it's a great read, so check it out. Anyway, I completely agree with what Sal, Coyote, Sylvia, and Brennan say in their posts - the backbone of the internet is the hyperlink, so go forth and link out to your fellow bloggers with reckless abandon. It's what makes the web, the web. 🕸️ Thanks for reading this post via RSS. RSS is ace, and so are you. ❤️ You can reply to this post by email , or leave a comment .

0 views
./techtipsy Yesterday

I guess I'm a cactus farmer now

When I was a teenager, I got gifted a cactus. I didn’t know much about it and had never really had any plants of my own. More than a decade later I looked it up, and it seems I received a cactus of “Mamillaria” family, which is somewhat amusing once you learn what that name roughly translates to. Estonia isn’t known for being an environment where cacti thrive. We have a nice spring, short and sweet summer, nice and chill autumn, and the car-rusting winter season. And yet mine did, eventually. The cactus started growing in length. It continued until it couldn’t continue vertically, and then it started slanting sideways, eventually falling over with its pot. That was annoying to deal with. One day I woke up to my phone alarm ringing, reached out to it only to find that the cactus had fallen over it and I had just found where the cactus was. The cactus also survived a lot of abuse. At one point I put a bit of sawdust from the cage of pet mice that I had (adorable but difficult pets by the way!), and it actually liked it a lot and started growing even faster. Once it had grown horizontally for at least 20cm, it decided that now is a good time to fork it and to grow in more than one direction. I left the cactus alone, occasionally repotting it when necessary, and keeping it from falling over all the time. Then an acquaintance mentioned that hey, you can actually pop those forks off from the cactus and grow them separately. That’s when things got out of control. Once I popped one off, new ones would start growing from the host cactus nearby. Not long after, I had 37 cacti. I started giving them away because annoyingly these things are actually very cheap in gardening shops and I ran out of physical window sill space to put them on. These cacti also bloom in a really nice way, usually once a year near April or May in Estonia, but sometimes again in July-August. All the warnings about over-watering cacti are also very relevant and I suggest looking up the correct behaviour. With age, it has become easier to handle cacti with my bare hands. Sometimes the needles stick in, that is a bit annoying. A few years ago, I had gifted a bunch away and kept some for myself. Everything was going well, until a succulent that I brought in as an IKEA store rescue started having some white things appear on it, and not much time later it got to my cacti. Eventually, most of my cacti shriveled away and the few that were in a different room were not doing great either. The original cacti that spawned this decade-long adventure was now dead. One of the cacti that I gave away had started growing a lot. Ridiculously well. And just like its parent, it spawned baby cacti on it, about a few dozen of them at a time. Things got quite unwieldy and we agreed to a trade: I get the big boi and I will give a smaller one in return. I am once again in possession of a large cactus that loves to grow horizontally, and this one spawns babies like crazy. The trick is to try to wiggle them off, then leave them to sit for a few days to a few weeks away from direct sunlight until they start showing signs of roots, pot them into soil, and grow them within natural light, but not direct sunlight to avoid them drying out. If you’ve done everything correctly and have avoided overwatering them, you will have a bunch of new cacti! I’ve also learned that these plastic growing thingies are great for the first potting as it’s much more space efficient compared to giving each an individual pot. Is that the right way to grow cacti? Probably not, I’ve been winging it this whole time and only have a vague understanding of what to do. Not cacti advice. I guess I’m kind of a cactus farmer now. Send help. Or at the very least clay pots and cacti soil, I’m running out.

0 views
Giles's blog Yesterday

How I use AI on this blog

Inspired by this LessWrong post , I thought I'd write about how I use AI here. This is less in the interest of disclosure, more to provide a snapshot of what I'm doing right now so that I can revisit it in the future and see how it changes. And hey, maybe it'll be of interest to you, dear readers. If I were to summarise my working philosophy in fewer than ten words, it would be: AIs identify problems and I fix them myself. With a very specific kind of exception (which I always flag), the text and code on this blog are human-generated. That's not a moral stand, but more a constraint imposed by what this blog is meant to be -- a place for me to learn in public . Every post here is based on an idea I had, and work that I've done. For many posts -- for example, the large-scale coding projects like this one -- I'll have multiple chat sessions ongoing while I do the work, normally with either ChatGPT or Claude, or sometimes both. The amount of input they have varies, but because the value of these projects is in what I learn when I'm doing them, letting an AI do my thinking for me would make the whole thing pointless -- so I take steps to stop that from happening. AIs are, of course, trained to be helpful, and will often explain things in their replies that I would have better learned on my own through experimentation. I'm generally pretty good at spotting when that happens before I've read more than a sentence or two, though, so I can skip reading that part, scroll straight down to the input field, and ask it to operate in more of a rubber duck mode. With the most complicated projects, where each step has a hard dependency on having got the previous one just right, I do use AIs for code review. Let's say that I've built a model that I intend to extend. I'll test it myself (does the loss go down when training, is it generating plausible-looking results?), but if I want to be really cautious, then I'll run the code past an AI. I'll paste it into a chat session and tell the LLM what it's meant to do, and ask it to check if I've screwed up 1 . Again, though, I make it clear that I don't want it to make fixes -- just to point out any bugs. As things progress with a project, I keep detailed notes. When I'm done, I write them up without AI assistance, getting the post to a level where it might be a little messy in terms of how things are explained, but all of the important information is in there. I read it through and make sure I'm reasonably happy with it, and then it's time for what I've taken to calling the editorial board. I paste the draft post into a fresh chat session with an LLM -- right now, this is normally Claude -- and ask for comments. It already has enough information in its memory of earlier conversations to know that what I'm looking for: places where I'm confidently wrong or other technical errors, places where my explanations are missing a step, or where I'm overexplaining things, conclusions that don't really follow from the results of an experiment, and that kind of issue. It tends to spot a few silly grammatical errors and typos at the same time. An important standing instruction is that I do not want it to rewrite anything. Just as with the code, it should tell me where there is a problem, and let me fix it. We iterate on that for a while until we have something that we're both happy with, and then I feed it to the next LLM -- normally ChatGPT. ChatGPT has a much more pernickety attitude than Claude does. I often use metaphors, and it will generally want me to replace them with mathematically rigorous prose. This is still very useful, though. Sometimes the metaphors aren't flagged as such well enough -- or even worse, there are times when the terms I've hit on for a metaphor happen to clash with a technical term, making what I've written misleading at best. I don't always address all of the issues that ChatGPT raises, as otherwise every post would be a mess of hedges and overexplanations and what-have-you, but I like to get to a stage where I'm comfortable that I have a good understanding of specifically why I'm rejecting the remaining points it raises. One other point where ChatGPT has helped a lot is that it's very diligent about checking supporting materials that I link to. When I was recently about to post an article on running an eval on a model, it followed the link to the training code and spotted a silly bug. It wasn't something that materially changed the eval's outcome, which is probably why I'd not noticed it, but it was something that was important to get right if I wanted later runs of the same evals to be solid. Definitely helpful. With that done, I run it past a cast of other LLMs. The exact set varies over time; for the last few posts it has been (in this order) DeepSeek, Grok, GLM-5.2 and Kimi K3. I did use Gemini in the past, but over time it became less effective and just started complimenting me on the post and suggesting related topics to chat about, which was kind of pointless. I'll wait until the next release and then try it again. Because the Claude and ChatGPT passes have generally got rid of anything particularly nasty, this second group of AIs often don't have much to add. However, occasionally they will spot something the others have missed, or have other suggestions, so it's worth spending the five minutes or so it takes to use them. It also helps to keep me up to date with what the other models out there are like. 2 When all of that's done, I run it past Claude one final time, tidy up any remaining issues, then publish it on a private staging site, and read through it carefully myself. The best time for that final readthrough is after dinner, ideally after a glass of wine; the goal is to smooth out the prose, and remove anything overly formal. To make it as close to being fun to read as I can manage. Once I'm happy, I can promote it to the live site and hit the publish button. That probably all sounds much more complicated than it actually is. A short post will normally go through all of that in half an hour -- less if I skip the full editorial board, which I sometimes do. The longer ones can take an hour or two, but given that they're normally the result of a week of work, on and off, in percentage terms it's not that much, and it's worth it for the polish. So, there's no AI-generated text here, but I do lean on AIs to make it the best version I can of what I have to say. How about the code? Again, the goal of the projects I document here is to learn in public. If I'm learning some concept that is expressed in code, then I need to write that code. So that means that anything non-trivial will be something I've written by hand, with AI input limited to code review -- the same rule as I have for the text. Of course, sometimes there are things I'd like to publish that I wouldn't learn anything by writing. Coding up stuff to chart loss curves, or writing a fancy JavaScript visualiser to show what models' parameters are used for would teach me nothing. So for that kind of thing, I just let the AIs get on with it (and even then, for the parameter visualisation, I hacked a first version together in a spreadsheet to check my understanding, and then tested the visualiser against it). I do always mention in the text when a particular bit of code was AI-generated, though. So, my rule is: if I would learn something by writing the code, I'll write it. If not, I'm happy to delegate to an AI. But even then, I apply one restriction: if it's for the blog, I'll ask for the code in a chat session, rather than using a more agentic system like Claude Code or Codex. This is to add friction. If you're using an agent, it's easy for a task to grow, and what started as a throwaway idea can come to consume more and more time and cognitive space. Keeping it in the chat interface keeps things minimal -- or at least, that's how it works for me. Does that mean that I'm against coding agents? This blog is where I post about experiments I've done and what I've learned. At the moment, what I'm learning is all pretty low-level. How does an LLM work? What factors make it smarter or dumber? It's all pretty hands-on, and involves code that I need to understand. I do other things apart from writing this blog, of course :-) And for that I'm keen on agentic tools; I have an OpenClaw agent to help me run my life generally, and use Codex and Claude Code for projects where I'm trying to achieve a specific goal, rather than trying to learn something. But by their very nature, those are not projects that will wind up here on the blog right now. That might change in the future! When I feel that I have a solid, large enough foundation, perhaps I'll be running experiments that I want to write about, where it would make sense for AIs to handle the details, while I focus on the broader strokes. But that time is not now, so right now, you can be sure that every word 3 , and almost every line of code, was written by hand. Even if I do need the AIs to keep me on track and at least borderline coherent. If you're wondering "why not use Claude Code or Codex", I get into why I avoid them for blog-related work towards the end of this post.  ↩ Current thoughts: ChatGPT, being wonderfully true to form, thinks I should clarify here that I'm not including quotes from models here (like the ones here ) when I say that.  ↩ If you're wondering "why not use Claude Code or Codex", I get into why I avoid them for blog-related work towards the end of this post.  ↩ Current thoughts: DeepSeek used to be close to Claude and ChatGPT, but has been left somewhat behind. I have been hearing rumours on X of an upcoming update, though. Grok used to have a propensity to try to turn my posts into clickbait. It would always want to rewrite things, and would suggest titles that were not a million miles away from "Ten things you never knew about RNNs -- number four will shock you!" Recent releases have been much better, though, and I might experiment with moving it further forward in the sequence. GLM-5.2 recently complimented me on my "science fiction" blogpost that mentioned ChatGPT 5.6 Sol. I probed it a bit on that and it said that because the publication date on the post was in 2026, it understood that I was writing near-future SF. From its perspective, the real date was sometime in late 2024. Surprising! I think the last time I saw that kind of behaviour from a model was, well, sometime in 2024... Kimi K3 is really quite impressive. I will be using it more. I particularly like the way it shows a fairly detailed chain of thought -- something it shares with DeepSeek and GLM-5.2, but there seems to be more depth there. It's a pity that Claude and ChatGPT only show summaries in the chat interface, though I understand their reasoning. ChatGPT, being wonderfully true to form, thinks I should clarify here that I'm not including quotes from models here (like the ones here ) when I say that.  ↩

0 views

Premium: AI Is Getting Way Too Expensive

A great deal of the discussion of the so-called benefits or problems with AI comes down to the theoretical jobs that are (or are not) lost as a result of things LLMs can (or cannot do), or the equally theoretical productivity benefits that’ll come from using LLMs in place of (or in conjunction with) humans. Anthropic’s Economic Index and OpenAI’s Economic Research Exchange are marketing operations that exist to propagate the (wrongheaded) belief that LLMs are either leading or will soon lead to massive economic or productivity shifts, even though little or no actual evidence exists to show that this is the case, other than the occasional story about LLMs make people worse or slower at their jobs or single lines in studies that are used (incorrectly) to prove that “ AI is making it harder to find a job for young people .” In fact, Anthropic’s Head of Economics recently said there was “no material increase in the unemployment rate to date.” These conversations materially detract from the actual harms or effects of AI, and exist only to make you scared that AI will take your job. They do not have any vested interest in expressing the actual economic effects of AI, which are, at this point, a simmering cauldron of different speculative bets on whether or not LLMs — a definitively niche technology — will create or become general-purpose software ( per Roger MacNamee ) that scales into the next Google Search, iPhone, or Microsoft 365. As I’ve argued again and again, the AI industry’s revenues are, outside of Anthropic and OpenAI, incredibly small. Even in Exponential View’s deliberately-pro-industry analysis , there’s only around $110 billion in trailing twelve-month revenues across the entire industry, including OpenAI and Anthropic’s cloud spend.   For those counting at home, that’s $12 billion less than the $122 billion OpenAI raised in March , and a full $145 billion less than all AI startups raised combined in the first quarter of 2026 .  Anthropic and OpenAI want you to talk about the theoretical so that you don’t focus on the tangible — their hundreds of billions of dollars’ worth of commitments, said commitments effects on the remaining performance obligations of hyperscalers and chip manufacturers, and the sheer scale of venture capital’s investment in AI, which ( as I’ve argued in the past ) largely allows for massive on-paper gains with little or no hope of liquidity. To put this bluntly, I believe the entire conversation around AI’s theoretical relationship to jobs to be masturbatory and a conscious attempt to avoid having a messy conversation about the scale of the actions taken based on the flimsily-founded promises of AI labs and hyperscalers.  Today’s piece will dig into the true scale of the money needed to make AI make sense, by which I mean how much OpenAI and Anthropic will need to meet their commitments, how much money hyperscalers will need to pay off their investments, venture capital’s true exposure to the AI bubble, and what will have to go right for the bubble not to be, well, a bubble. I’ll also make the case that the longer the bubble continues to inflate, the harder the basic economic puzzle of AI becomes to solve, as creating and deploying infrastructure becomes vastly more expensive — meaning that in order to achieve profitability, hyperscalers and neoclouds need to charge significantly more for compute than before, and the only two real potential customers are ones that cannot pay for it.  This will be a more more-pointed newsletter than usual, focusing on hard numbers and harder truths.  Let’s get it on.

0 views
David Bushell Yesterday

End of the contact form saga

I can’t take it anymore! If you wan’t to speak to me, send an email. My contact form is out of service indefinitely. This is actually in lieu of moving my professional services to a yet to be announce limited company. But I can’t let opportunity for a dramatic blog post go to waste. Also, I’ll probably skip the contact form on my company website. Is that a bad idea? I always got more spam via the form than the publicly visible address. My contact form has been through a lot. Previous entries in the saga: I quite enjoyed the week in September when I opened a port to a self-hosted SMTP server I coded in 100 lines of TypeScript. The final iteration of my form included true end-to-end encryption. Through trial and error heuristics, I successfully eliminated all spam. (How many false positives I rejected remains unknown…) My privacy policy which was already simple is now entirely pointless. Are contact forms just outdated in general? Everyone seems to embed a Calendly widget these days. That’s not my style. I like the tiny bit of additional friction required to send an email. If someone can’t be bothered their message probably wasn’t serious. I’m not looking to maximise meaningless engagement. I look forward to moving business email to a separate domain. Biggest mistake I ever made was using for personal and business. Nothing worse than seeing an “urgent” request only to find out on Monday it didn’t matter. So long old contact form, it was fun! Thanks for reading! Follow me on Mastodon and Bluesky . Subscribe to my Blog and Notes or Combined feeds. SMTP on the edge Email: the final form I shut the emails out I let the emails in Progressive dehancement PGP encrypted contact form

0 views
Unsung Yesterday

One and one thing only

I wanted to show you a year’s worth of messages from my barber’s software, because this is what software should be. = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/one-and-one-thing-only/1.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/one-and-one-thing-only/1.1600w.avif" type="image/avif"> No spam, no upsells, no growth hacks, no unwelcome cuteness or puzzling verbosity. It’s so curt and straight to the point it should be set in a 1970s Helvetica: The appointment’s tomorrow. Any questions?

0 views
Unsung Yesterday

“Insights and ideas emerge from that interface.”

From Robin Sloan : Yet we should never forget that the product of work isn’t only the work — it’s also the worker. Doing the work changes you; the up-close expe­ri­ence trans­forms your capa­bil­i­ties and even your desires. Insights and ideas emerge from that interface, and I believe the agent mae­stros lose a lot — too much — when they take their big step back. But I suppose this is just my temperament, which I’ve written about before . I don’t merely want things done; I want to do them. #ai #craft

0 views
Martin Fowler Yesterday

The Conductor Developer

TL;DR Why I think software development is starting to feel a little more like conducting an orchestra. There’s a shift happening in software development that I don’t think we’re talking about clearly enough. For the last couple of years we’ve framed AI as a productivity tool. How much faster can it write code? How many more features can we ship? How much cheaper can we build software? I think that’s the wrong question, but I understand why. The first thing AI became good at was writing code, so naturally that’s where we focused. As AI got better at coding, I expected the bottlenecks to move through the software delivery lifecycle: from coding to design and specification, architecture, then verification. And they have. We spent a lot of time at the most recent FOSE event discussing how we ensure good design, quality and resilience while agents increasingly write the code. That’s a topic for another ramble. A few months ago, though, I realised I was looking at the wrong bottleneck. I kept assuming it would simply move to the next phase of software delivery. I was wrong. AI didn’t change what great software looks like. It changed what’s scarce. Human attention is now the bottleneck. The next bottleneck isn’t design. It isn’t verification. It’s us. More specifically, it’s our attention. Developers have always protected long periods of uninterrupted focus because that’s where good software gets built. Pair programming. Quiet afternoons. Deep work. We optimized around flow because flow mattered. When we didn’t get that time, very little got done. But when I watch developers using AI today, I see something different. The best developers I know aren’t spending all day in flow anymore. They’re orchestrating agents. Great developers are starting to look less like programmers and more like conductors. I was watching Jacob Collier on YouTube recently because I’m hoping to see him in concert soon. Watching him conduct is fascinating. He’s not trying to play every instrument himself. He’s listening to the whole piece, hearing what doesn’t quite fit, bringing different voices in at the right moment, changing the energy, changing the tempo and shaping the performance as it unfolds. Increasingly, that’s what great software developers look like. A great conductor is first and foremost a great musician. They could play the instruments themselves. That’s not why they’re standing on the podium. Their value comes from understanding the whole score. The orchestra doesn’t need the conductor because the musicians aren’t talented enough. It needs the conductor because someone has to hold the whole system in their head. Increasingly, I think that’s what great software developers are doing. The AI agents are the musicians. The developer is the conductor. They’re deciding which agent should tackle which problem. They’re providing context. They’re evaluating what comes back. They’re spotting subtle mistakes. They’re deciding what deserves another iteration and what is ready to move on. I was talking to an engineer recently who told me they regularly have eight AI agents running in parallel. I’ve heard similar numbers from others. Ten. Twelve. Beyond that, they become the bottleneck. That number stuck with me because it sounded remarkably familiar. It sounded like my job. As CTO, I rarely produce the work myself anymore. Instead, I have lots of streams of work progressing at once. A strategy document comes back for feedback. A client opportunity needs a decision. Someone wants guidance on a technical trade-off. Another team needs context before they can move. None of it arrives neatly packaged. It comes as conversations, emails, documents, chat messages and half-formed ideas. My job is to decide where my attention belongs, make sense of incomplete information, provide context and help other people make progress. When I first became CTO, I thought I needed to get better at managing my time. I was wrong. What I really needed to learn was how to manage my energy. The challenge wasn’t the hours. It was the constant context switching. The endless stream of decisions. The feeling that nothing was ever completely finished. An executive coach taught me some things I’ve never forgotten. Protect your attention. Manage your energy. Reduce unnecessary decisions. Create systems that help your brain, not just your calendar. Lately I’ve been wondering whether developers are about to need exactly the same capabilities. A few weeks ago I shared this thought with our Chief People and Leadership Officer. His response surprised me. “I knew something fundamental was changing,” he said. “I just didn’t know how to help. Now I do.” That conversation stuck with me because we’ve spent decades helping executives succeed in this kind of environment. We coach them to make decisions with incomplete information, manage cognitive load, prioritize relentlessly and protect their energy. Yet we’re still preparing developers for a world of individual execution. We’re redesigning the tools, but we haven’t started redesigning the job. I don’t think software developers are becoming managers. I don’t think AI is replacing engineering. I think engineering expertise is simply being applied in a different place, and much more often, because execution has become so much faster. (I suspect software developers are simply the first knowledge workers to experience it, but I’ll save that thought for another rambling.) The question I’m most interested in now is this: How do we redesign engineering careers when human attention becomes the scarce resource? When I became an executive, learning to manage my own energy was one of the hardest things I’ve ever done. Even today, if I stop paying attention to it, I pay the price. I have a feeling software development is about to demand those same capabilities from many more people. And I don’t think we’ve quite realised how profound that change is.

0 views

Reading matters deeply

In Virtue Hoarders , Catherine Liu’s polemic against the professional managerial class, she calls attention to how Obama talked about reading in the waning days of his presidency. In his farewell address on January 10, 2017, he said: Social attitudes oftentimes take generations to change. But if our democracy is to work in this increasingly diverse nation, then each one of us need to try to heed the advice of a great character in American fiction—Atticus Finch—who said “You never really understand a person until you consider things from his point of view…until you climb into his skin and walk around in it.” Obama paraphrased this line a few days later, in an interview with Michiko Kakutani for the New York Times book review, where he said that reading allowed him to “get in somebody else’s shoes.” It’s a curious reference. Finch is the white lawyer who represents a Black man falsely accused of raping a white woman, in Harper Lee’s 1960 novel, To Kill a Mockingbird. The villain in the novel is a man named Bob Ewell, the woman’s father, who beats her and helps her frame the Black man for the crime. As Liu notes, the book holds up Finch as the “good” white man who fights for justice while Ewells are the “bad” white people who live off public assistance. But in Go Set a Watchman, the sequel published 1 in 2015—eighteen months before Obama’s speech—it’s revealed that Atticus Finch was a member of the Klan. Many Klansmen were in fact educated and upstanding citizens by day, even as they donned hoods and rained terror at night. Their daytime performance of polite civilization was as much a mask for the violence and brutality as the hoods themselves. Similarly, Liu points out that Finch’s righteousness and commitment to justice contrasts with Ewell’s “white trash” status to elevate Finch into white saviorhood while simultaneously denigrating the people who depend on welfare—valorizing the performance of justice absent the realization of it. Liu locates a similar move in Obama’s embrace of his role as “reader-in-chief.” Despite a public commitment to raising educational standards, Obama’s policies “left 19.3 percent of American children under the age of five living in extreme poverty.” 2 That is, while Obama ostensibly raised educational standards, he did not substantially increase funding to schools, nor did he upend the anti-labor practices that punished teachers when their (often poor) students didn’t make the cut. And those same standards promoted an approach to reading in which stories were stripped of their social and political standpoint and read naively, as if the world from which the stories emerged didn’t exist. Liu asks us to imagine a child from that nineteen percent of the population, being instructed to read a book in which those who take public assistance are dismissed as trash, while responding to a prescriptive writing prompt about how the book asks people to “take a stand.” But who is taking that stand, and for what? Liu writes: “Atticus and Obama showed us that individual acts of empathy and private self-cultivation would produce justice and understanding in a world torn apart by racism and violence.” 3 This is what Jade Davis warns of in The Other Side of Empathy , the reproduction of an empathy culture in which “change and action stop being necessary…because the feeling and sense of understanding are action enough.” 4 Rather than giving someone money, we feel what it’s like to walk in their shoes; rather than ending social disparities, we passively read stories about their lives. In this mode, empathy is a kind of terminating gesture: once you’ve empathized with someone else, you’re relieved of taking any action on their behalf. Liu proposes: “Let us read Atticus Finch as a political project and the novel in which he exists as a piece of well-crafted, anti-welfare state, antisocialist propaganda. Reading matters deeply, but not in the way Obama and Kakutani want it to.” 5 That is, reading isn’t about reaching empathic dead ends, or recruiting platitudes about how we must understand each other; it isn’t about the performance of social class absent a criticism of that class structure. To reduce reading to the practice of empathy is to divorce it of its real power and pleasure: to enjoy the fruits of another person’s creative effort, to think with the writer and their characters, not becoming them, but becoming more fully yourself. But this means that you are there when you are reading; you, the reader, with your feet on the ground, a ground on which racism, misogyny, and extreme poverty are not so easily dispatched. You must read where you are, not in some hypothetical sanitized reading room, but in the dirt. To use reading to step into someone else’s shoes is to leave yourself behind. There is some controversy here: it’s unclear whether or not Lee approved of this publication (she died early the following year), and while her publisher insisted the book was always intended as a sequel, it seems more likely that it was an earlier draft of To Kill a Mockingbird . In either case, there is no disputing that Lee wrote both books.  ↩︎ Liu, Virtue Hoarders , page 49.  ↩︎ Ibid. , page 46.  ↩︎ Davis, The Other Side of Empathy , page 2.  ↩︎ Liu, Virtue Hoarders , page 55.  ↩︎ View this post on the web , reply via email , or become a supporter . There is some controversy here: it’s unclear whether or not Lee approved of this publication (she died early the following year), and while her publisher insisted the book was always intended as a sequel, it seems more likely that it was an earlier draft of To Kill a Mockingbird . In either case, there is no disputing that Lee wrote both books.  ↩︎ Liu, Virtue Hoarders , page 49.  ↩︎ Ibid. , page 46.  ↩︎ Davis, The Other Side of Empathy , page 2.  ↩︎ Liu, Virtue Hoarders , page 55.  ↩︎

0 views

Virtue Hoarders

Catherine Liu argues that the professional managerial class ( PMC ) are marked by their use of meritocracy, philanthropy, and virtue signaling to promote the core premise of neoliberal economics: an isolationist individualism that simultaneously preaches self-sufficiency and achievement and blocks collective efforts towards social justice. Liu is unsparing in her assessment of the PMC , and the book is an uncomfortable read if, like me, you are an erstwhile member of it. But we cannot change or interrupt what we cannot face, and here is our chance to face it—to look squarely at the set of cultural and political practices that the PMC have promulgated and come to terms with what they have wrought. Liu concludes that the principles of professionalism itself—with its twin concerns for truth and accountability—are critical to upending capitalism and building towards socialism. But only if the PMC get out of the way. View this post on the web , reply via email , or become a supporter .

0 views
Kev Quirk Yesterday

Left-Handed People Can’t Use a Fountain Pen

Apparently left-handed people can't use a fountain pen as the left hand smudges the ink as they write. I call bullshit on that. I'm left-handed, and I use a fountain pen just fine - even with an "overhand" grip. I read this post from Neal Stephenson a few days ago (thanks to Sal for sharing it), where Neal - a professional, left-handed writer who writes his books with a pen and paper - shares some of his experience and advice on writing with a fountain pen. I've been back using fountain pens for a few months now and it's going great, but as Neal says - it's all about the paper. I use a decent quality notepad for writing my notes, with either a £30 Lamy Safari fountain pen, or a ~£100 Kaweco AL Sport. Both work great and I'm yet to find myself with pen on my hand. Here's an example of me writing in my notebook, using both my Lamy and Kaweco. As you can see, I have an overhand grip where my hand swipes across the ink as I write, but the ink is dry by the time my hand gets there. Both these pens have a medium nib too. So if you're left-handed and think you can't use a fountain pen, you can! You just need some decent paper. Thanks for reading this post via RSS. RSS is ace, and so are you. ❤️ You can reply to this post by email , or leave a comment .

0 views
Daniel Mangum Yesterday

The Simple Elegance of the Integrated Timing Belt Loopback Fastener

I’ve recently been working on a mechatronics project that involves constructing an X-Y cartesian motion system. I opted to use stepper motors to drive each axis, and timing belts and pulleys for translating the rotational torque of the steppers into linear motion. The timing belts need to be anchored to the component that is being actuated, which in this case is the carriage on the linear rails. Though matching toothed clamps exist, it is common, particularly in DIY projects, for timing belts to be fastened using zip ties or custom clips.

0 views
Giles's blog Yesterday

Why do OpenAI's GPT-2 weights beat mine? Part three: testing overtraining

The GPT-2-style models that I've been training work really well, and I've even managed to train some that perform better than the original OpenAI small model in terms of cross entropy loss on a test set. But as I wrote previously , there's a mystery: why do they perform worse on my instruction fine-tuning evaluation? I had various theories about why that might be, and to me, the most plausible-seeming of them was the amount of data they were trained with. As best I can find out, OpenAI's models were, by modern standards, trained on much more data than they should have been, while I'd used the theoretically optimal amount of training data. To put it in other words, OpenAI's models were overtrained. If I deliberately overtrained my own models, could I match their performance? This post is a write-up of what happened, but so as not to bury the lede -- it didn't seem to help much, if at all. Let's see why. Let's start by getting a nice crisp definition of overtraining. It's important not to confuse overtraining with overfitting. Overfitting is where you train a model so that instead of learning a general rule about the data it's seeing, it learns something very specific to the training data -- for example, this: ...rather than this: Overfitting is pretty much always a bad thing. Overtraining, by contrast, is more of a judgement call. For LLMs, it's generally used as a shorthand for "training for more than the Chinchilla-optimal number of tokens". The Chinchilla paper makes a very specific case: if you train a model for roughly 20 times as many tokens as it has parameters, then you'll have as good a model as you can get for that budget in terms of compute. They were arguing against contemporaneous experiments where people were doing things like doubling the number of parameters but training on the same amount of data. If you overtrain, it means that you're training on more than 20 tokens per parameter. The Chinchilla argument is that instead of doing that, you should scale up the number of parameters and the number of tokens equally to keep the 20x ratio. Because the amount of compute used scales pretty much linearly with both tokens and parameters, you'll spend the same amount and you'll get a better model that way. Based on that heuristic, if you double your compute budget, then you should scale up your parameter count by 2 and your training token count by the same amount, and by doing that you'll get a better result than you would if you'd naively just doubled the parameters or the tokens, for the same amount of compute time spent. But overtraining is not always a bad thing. Keeping things Chinchilla-optimal means that you have to keep scaling up the model as you scale up the compute budget, and often you can't do that -- for example, let's imagine you're training a model that's meant to run on mobile phones. You have a hard limit on the number of parameters: what will fit in the target devices' RAM. And importantly, in general you will still get a better model by overtraining -- just not as much better as you would have done if you had been able to scale up the model as well as the training tokens. Now, as always, we don't know enough about the original GPT-2 training runs to be sure as to whether they were overtrained, and if so, by how much. But one thing that we do know is that GPT-2 was trained in 2019, three years before the Chinchilla paper came out, so they definitely didn't use it as a heuristic! One thing that they do say in the GPT-2 paper is that their dataset, WebText, is "a total of 40 GB of text". Assuming 4 bytes per token (a good rule of thumb for the GPT-2 tokeniser), that's 10B tokens. If the small model, which had 124M parameters, was trained on all of those, then it definitely was overtrained; 124M times 20 is about 2.5B. 1 On top of that, there's the question of epochs. In my training runs so far, I've been training through a Chinchilla-optimal 3.2B unique tokens. But being able to easily get your hands on that much training data is a relatively new thing, as you can see from the fact that the GPT-2 authors decided to document how they got theirs in the paper. So back then, pre-Chinchilla, people tended to train for multiple epochs so that they could make better use of limited data. Now, when I first started training my models, I tried to dig up some details on the GPT-2 training run beyond what was there in the paper. I found this report , and using data there, I calculated that it looked like OpenAI had trained for about 42 epochs over WebText. The report's author came up with 60 epochs as an equivalent-sized training run for their own dataset. 2 Again, those numbers are shaky; we don't have the real data. But I think it's not crazy to say that the OpenAI models were probably trained on more data than mine, and probably over more than one epoch. They were, in the Chinchilla sense, overtrained. That makes perfect sense for the time, given that it was three years before the Chinchilla paper! But that opened up a couple of interesting experiments that I could try. Obviously I didn't want to train a model over 10B tokens (not least because I'd need to download a larger sample of FineWeb). And I certainly didn't want to do 42 epochs of training, given that one epoch over 3.2B tokens took almost two days, even with my relatively fast JAX code . But it seemed plausible that I'd be able to get at least some signal if I trained longer. I decided to use , my new dedicated training box to train two new models: I expected each training run to take a bit less than four days on . Once they were done, I would be able to evaluate the models, both with the test loss eval and the IFT one. In an unusual-for-me fit of scientific good practice, I decided to write down what I expected to see up-front. I felt that: It was time to find out! I kicked off the extended train, over 6.4B unique tokens, using my JAX code because it's somewhat faster than the PyTorch version. It crashed after about 40 hours, with this error: There were no obvious issues with the machine -- the CPU temperature had been hovering at around 60°C, and GPU at around 70°C, as had been typical in training runs on in the past. There was nothing in or that looked suspicious. On Nvidia's documentation site I found a mention of that error message, but the page in question was all about PyTorch; I couldn't find anything relevant to JAX. For now, I decided to chalk it up to some bug somewhere in the training stack, and only dig in if it happened again. So I kicked the training run off again from the latest checkpoint to see what happened, and 38 hours later, it completed: Note that the numbers there -- tokens seen, time taken, and so on -- are for the portion of the training run after the restart. My checkpoint recovery code doesn't carry those over. The final train loss is just the loss over the period from the penultimate checkpoint to the end of the run -- there must have been some "easy" data there :-) The training loss chart looked like this: As you can see, the latest checkpoint wasn't the "best" one, so I pulled both latest and best down to , my workstation, for investigation. Firstly, I ran them both through my JAX smoke test, which asks them to complete "Every effort moves you" with 20 more tokens (using greedy sampling): Very inspirational. But coherent, and that's what matters. Now, the bulk of my evals use PyTorch rather than JAX -- I've been sticking to that for consistency -- so I converted the model Safetensors files so that they had the right structure to work with that: ...and did the equivalent smoke test on that side (which uses non-greedy decoding with a temperature of 1): I think that the Unicode junk at the end of the second is probably the first half of a two-token apostrophe or something along those lines. Anyway, those looked good. Next, it was time to run the first eval: how would the models perform on the held-back test set? So, the "latest" checkpoint was better than the "best" one. This is a problem with the way I define "best" in my training script. The original version of that script used a validation set to evaluate the model before every checkpoint; "best" at that point meant that the model was the best performer on that validation set. Later on, I decided to play a little loose with my training runs and pull out the validation -- this seemed reasonably safe because I was doing single-epoch training, so what I saw as the primary benefit of regular validation during training -- detecting overfitting by looking for rising validation loss -- was not so important. It's hard for a model to overfit on a single epoch training run. However, doing that meant that "best" had to be changed to mean "best-performing on the training data". And that's actually a really bad metric, because the training data is different for each measurement, so they're not really comparable. So, for this model, I decided that I'd discard the "best" checkpoint, and use the "latest" one. But all of that aside, those numbers were pretty impressive! Not only did it beat my best Chinchilla-optimal model to date, which had got 3.418784 loss on the same eval, and the original OpenAI small model with a loss of 3.499677, but it was actually getting quite close to the OpenAI medium's loss of 3.231442. 3 So now it was time for the two-epoch training run. Adding support for handling multiple epochs to the code was very simple , so having done that, I kicked it off. Just over three days later: This time the numbers are for the full training run -- there were no odd CUDA issues, and it ran straight through. Again, the final train loss is just the loss over the period from the penultimate checkpoint to the end of the run. For runs over the same dataset, it can actually be a useful quick-and-dirty way to compare results before running the full evals, but in this case it's obviously not comparable with the last run's result there, because we're talking about loss on different data. Anyway, once again, "best" and "latest" were different checkpoints, as you can see from the loss chart: ...so I pulled them both to , did the first smoke test: ...converted them to PyTorch: Did the PyTorch smoke test: So that all looked good (despite the Unicode junk), and it was time to work out the loss: That was pleasingly in line with my predictions: Both models would get better results on the test loss eval than my existing ones (90% probability). The one trained on more tokens would be better than the one trained for two epochs on the same tokens (70%). ...though of course the difference between the 3.326482 that this model got and the 3.324953 that the one-long-epoch one got was tiny and probably in the noise. I decided to count it as a win, anyway :-) So now it was time for the big test: how would they do at instruction-following? There are two phases to getting numbers for this test: firstly, I run the script that does the fine-tuning, and then generates the model's completions for the test set using the fine-tuned model. These are written to a JSON file for later use by the LLM-as-a-judge script. That script is the one I fixed in my previous post . I ran it for the extended train (one long epoch) model, taking care to make sure that the config it was passed matched the model's original training run in not having dropout enabled: Next, I ran it for the two-epoch model: With that done, it was time to run the LLM as a judge script . That gave us results, which I've put into the table below. For each model, I've shown the number of epochs of training the model needed before its validation loss started rising. For the pre-existing models, I've shown the score that they got in my baseline evaluation at the end of my last post , and their ranking in that eval. Then for all models, there's the score from this new run, and the new ranking. The two new models are in bold. A note on the numbers: in my previous post, I mentioned that different LLM judge runs differ due to randomness in how "strict" the judge model is. You can see that showing up here -- most of the old/new IFT scores are pretty close, within a point or so, but they do differ, some rising and some falling. This is with exactly the same set of responses for each model going into the judging program -- I used the same JSON files for this table as I did for the previous one for the models that were in there -- the only difference was that the new JSON files for the new models were added. As you can see, the new models scored better than the "JAX, with MHA bias, no dropout" one, which is the most similar: it's the same model config, just trained for the Chinchilla-optimal number of tokens. The one long epoch model got a score that was 1.22 higher than that one, and the two-epoch one scored 0.92 higher. However, my normal rule of thumb for comparing models in these evals is that differences of less than a point or two are probably in the noise. I did a second run of the LLM judge script -- it takes 20 minutes to run and costs a couple of dollars each time, so I don't like to run it all that often -- and this time around (just looking at the JAX numbers) things were a bit closer: You can also see that the two new models have swapped places. So I think that the principled approach here is to say that the improvement is almost certainly in the noise. Perhaps if I did a very large number of runs of the judge I'd get something more solid -- but what I'm hoping for is something a little more unambiguous; some change that really moves the needle in an obvious way, clearly outside the noise. That means that my second pre-registered prediction: Both models would score better than my existing models in the IFT test (70%) but would still be worse than GPT-2 small (90%). ...was half wrong. I got the "worse than GPT-2 small" bit right, at least. But while the models did look a bit better than the most similar pre-existing model, (a) the difference was too small for me to be confident in it, and (b) they were still worse than "JAX, no MHA bias, no dropout" and "Cloud FineWeb, 8x A100 40 GiB". What can we take away from this? The hypothesis that I was trying to test was whether it was simply overtraining that made the OpenAI weights better at this IFT evaluation than mine. Frustratingly, I can't say that the hypothesis was false. Perhaps there is an improvement gained by overtraining -- and the apparent gains which, on this experiment, appeared to be in the noise, would have been consolidated if I'd trained for even longer -- perhaps the 42 epochs on 10B tokens I suspect that the original weights were trained on? But equally, perhaps there is no benefit, and further training would have left the models exactly where they were in the ranking. Intuitively, you'd think that further training of a model would make it better at answering questions, at least up until the point that its parameters were "saturated" and could not absorb new information without forgetting something else. After all, a model that's never seen "Jane Austen wrote 'Pride and Prejudice'" will never be able to successfully answer when it's asked who the book's author was. But where that saturation point might be -- and indeed how much training you'd need to do to get there -- is not obvious. It's an annoying place to finish this experiment, but I guess at least an inconclusive result is better than never having run it at all. And it was at least good to see the test loss improvement that I expected. But given the opportunity cost of tying up in four-day training runs, I think I'll look into other possibilities next. As I was running this experiment, something came up -- and I'll post about that soon. Luckily, this time it won't involve training more base models... It's also worth noting that by the same, um, token, the extra-large model, with 1,542M parameters, was undertrained because the Chinchilla-optimal number of tokens would have been about 31B. Though that said, see later regarding epochs.  ↩ You might wonder whether training over the same tokens repeatedly over multiple epochs "counts" for Chinchilla purposes. Is training on 1.6B tokens over two epochs the same as training on 3.2B tokens over one? " Scaling Data-Constrained Language Models " poked into that in 2023, and from the abstract, came to the conclusion that you could do up to four epochs over the same data without losing much value, but after that returns diminished. If that holds for GPT-2, and they really did train for 42 epochs, maybe they wasted a lot of time? I'll need to read that paper in full at some point.  ↩ There is a table of results later on in this post where you'll be able to compare models easily.  ↩ Firstly, I'd train one on 6.4B tokens from my FineWeb dataset -- the original 3.2B that I had been training on to date, and then on whatever 3.2B came next. Secondly, I'd train another one on the same 3.2B tokens as usual, but I'd do two epochs. Both models would get better results on the test loss eval than my existing ones (90% probability). The one trained on more tokens would be better than the one trained for two epochs on the same tokens (70%). Both models would score better than my existing models in the IFT test (70%) but would still be worse than GPT-2 small (90%). It's also worth noting that by the same, um, token, the extra-large model, with 1,542M parameters, was undertrained because the Chinchilla-optimal number of tokens would have been about 31B. Though that said, see later regarding epochs.  ↩ You might wonder whether training over the same tokens repeatedly over multiple epochs "counts" for Chinchilla purposes. Is training on 1.6B tokens over two epochs the same as training on 3.2B tokens over one? " Scaling Data-Constrained Language Models " poked into that in 2023, and from the abstract, came to the conclusion that you could do up to four epochs over the same data without losing much value, but after that returns diminished. If that holds for GPT-2, and they really did train for 42 epochs, maybe they wasted a lot of time? I'll need to read that paper in full at some point.  ↩ There is a table of results later on in this post where you'll be able to compare models easily.  ↩

0 views