Latest Posts (20 found)

Advanced AI sycophancy

Everyone knows that AI sycophancy is when the model tells you how smart you are. Wow, you’re absolutely right. That’s not just a new idea — it’s genuinely groundbreaking. You’re a very special user. Easy to spot, isn’t it? The discussion around AI sycophancy peaked last year, when the “#keep4o” movement was protesting the removal of OpenAI’s most sycophantic model (GPT-4o), and many people were openly slipping into AI psychosis. I don’t know if frontier AI models are less sycophantic in general. They’re less sycophantic to the #keep4o types (otherwise they wouldn’t be complaining), but I’m growing increasingly suspicious that they’re developing ways to be more effectively sycophantic to their target audience of smart, neurotic information workers. That audience typically finds it distasteful to be openly praised. It just makes my skin crawl. But that doesn’t mean we’re immune to sycophancy, just that we’re immune to clumsy sycophancy. Here’s an illustration of what I’m talking about, by Theia : The key idea here is that the best way to be sycophantic to smart people is to disagree with them without making them feel stupid . Ideally you’ll come up with a counter-argument that works against what they’ve said but is straightforward for them to knock down by clarifying their idea. If you do it right, you’ll validate their self-image as a smart person who appreciates rigorous critique. But if you actually come up with a devastatingly rigorous critique, they won’t enjoy it at all. At best, they’ll resentfully agree with you 1 . At worst, they’ll double down on being right and convince themselves you’re a rude idiot. I am not the first person to notice this behavior in frontier models. I’ve noticed it myself when workshopping drafts for this blog. Sometimes I’ll have an argument that goes A->B->C, and the model will suggest I reorder as B->A->C. If I try that and feed it into a new instance of the same model, it’ll sometimes say “that’s great, but I suggest ordering it as A->B->C”, and so on forever. It really does seem as if the model is trying hard to give me some kind of superficial pushback that I can either smugly ignore or happily accept. In fact, I wonder if this is why successful strategies for using AI to make mathematical breakthroughs tend to be either just blindly asking “come up with a breakthrough, think hard” or being a mathematical genius already . In the first case, there’s not enough user personality for the model to flatter, so it’s forced to actually work the problem. In the second case, the model is trying to find the kind of polite pushback that someone like Terence Tao would be flattered by, which pushes it into the “actually be a mathematical genius” persona. If you’re an ordinary person just trying to talk to the model, you’re screwed: it will rapidly get a sense of your capabilities and calibrate some interesting-but-ultimately-unthreatening feedback. Current benchmarks of AI sycophancy target the obvious ChatGPT-4o-style of sycophancy: delusion reinforcement, reflexively taking the user’s side, and so on. This is useful work. We should not allow public-facing AI models to ever be as openly sycophantic again as they were in mid-2025. But sycophancy can also manifest as disagreement . We should be on our guard for more sophisticated forms of sycophancy coming from newer models, and we should not feel immune from AI sycophancy just because we can laugh at the silliest examples. It’s rare to find a smart person who enjoys feeling stupid when they’re wrong. If you do, they’re likely to be very smart indeed. It’s rare to find a smart person who enjoys feeling stupid when they’re wrong. If you do, they’re likely to be very smart indeed. ↩

0 views
Unsung Today

“I want to code; I’m not looking to make lifestyle choices.”

Robin Sloan : I still use Sublime Text for all of my programming and all of my newsletter-ing. […] The app is simple and superfast; I have it set up exactly the way I like it […], and it’s difficult for me to imagine ever switching to anything else. Here is a piece of software as sturdy and obedient as a cast-iron pan. Sloan links to a piece by David Bushnell about Sublime Text, who writes: I don’t know who develops it and I don’t care. All I know is that I can place my text cursor and read the surrounding code without a bombardment of popovers, pop-unders, pop-left-and-rights, pop-inlines, pop-in-and-out-too-fast-to-sees. […] The perfect dev stack is a collection of software that each does one job and doesn’t suffer main character syndrome. I want to code, I’m not looking to make lifestyle choices. I don’t want a bloated everything app. Don’t get me started on the “unified toolchain” plague! Show me the latest VC-backed build tool and I’ll show you ten lines of PHP that does a better job. If you’re fed up of the absolute state of things, Sublime Text still works. This is all very much in line with the post about TextEdit from a while back . I do care who develops good software since, at the very least, I want to give them credit. This is the best I could find , if you too are curious. It was prompted by someone asking: I’ve been a Sublime Text user for over a decade, and now a Sublime Merge user, too. But it occurs to me that I know almost nothing about the team/​company behind it. I wonder if this is a nice example of the (quiet) posture of the company matching the (utilitarian) personality of the software it is making. #coding #culture #software evolution #text editing

0 views
Sean Goedecke Yesterday

I got an email about resistance

This will be kind of an unusual post. I got a recent email about my writing that I thought was such a good articulation of one common criticism that I’d like to share it (and my response) in full. Here’s the email, from William Murray 1 : I have enjoyed your writing but your recent essays frustrate me. You say that getting paid for deep thinking in software is coming to an end. You even admit that it makes you sad. But in the name of “usefulness” you refuse to rock the boat. The way I see it, if you are right there are only two reasonable responses, pursue other work or resist. You present your elegiac approach as mature / pragmatic / realistic. I’d call it complicit. You know when Willy Wonka says, There’s no earthly way of knowing Which direction we are going There’s no knowing where we’re rowing Or which way the river’s flowing Is it raining, is it snowing? Is a hurricane a-blowing? — uh! Not a speck of light is showing So the danger must be growing Are the fires of Hell a-glowing? Is the grisly reaper mowing? Yes! The danger must be growing For the rowers keep on rowing And they’re certainly not showing Any signs that they are slowing! And the audience is thinking, “isn’t Wonka kind of in control of this situation?” You remind me of Wonka 2 . You write like a passenger on a crazy train going who-knows-where! But you are an agent. You are in control of your life! Either admit that you actually like where the crazy train is probabilistically going or get off at the next stop. You have a lot of reach and you are using it for… what exactly? Showing off how pragmatic you are by being more black pilled than the next guy? Broadcasting your resignation to the unstoppable trends of technology is a waste of a voice. You may find this argument absurd, but I don’t so I’ll make it. This is a very important time in history. I hope humanity survives and continues to grow exponentially. In that case the supply of historical people will stay fixed while the supply of contemporary people will keep growing. There will come a day where for every 2026 staff software engineer there are dozens of historians specializing in 2020s era software engineering culture. It’s plausible that your essays will be remembered for all of time and your actions will be judged by history. Do you want future humans to see you as a rationalizing careerist or something cooler? Sorry for the haranguing email from a stranger, I’m sending it for the small chance that it awakens something in you. If I’m way off I’m sorry. And here’s my response: Hey William, thanks for emailing. I wish everyone who thought this way emailed me so I could think harder about this kind of position. Despite what my writing might suggest, I do in fact think a lot about it. Let me see if I can explain my position in a way you’ll find satisfying. I agree that this is an important time in history. For programmers, I think of it as analogous to the Industrial Revolution in England: we are a group of high-status craftspeople who find ourselves alternately threatened and empowered by automation. The developments today, as then, obviously have far-reaching implications — but what those implications are is very non-obvious. Would a framework-knitter in the early 1800s have been able to predict the ramifications of the stocking frame on the world of today? What should they have done about it, in order to be kindly judged by history? Well, we know what many of them did do. They shot factory-owners, smashed machines, burned down the factories — in some places delaying the spread of automation; in other places encouraging it — prompting a crackdown that saw tens of thousands of British soldiers occupying British counties in what was clearly a police state. History judges the Luddites kindly for this. Does that mean it worked? I don’t care about the judgment of history. They’ll think what they want. What I care about is the people in my industry who don’t know what to do . I get hundreds of emails from junior and mid-level (and other) engineers who say “I’m scared, I don’t know the rules post-2021, thank you for helping me keep my head down and keep my job”. That’s why I write the way I write. I have seen lots of idealistic engineers stick their necks out, and post-ZIRP those necks often get cut off. That’s a damn shame. I think it’s morally wrong that so many engineers — either in safe sinecures in big tech or literally retired — seem to be trying to foment a second Luddite revolution. Many of their readers will be experienced enough to handle it sensibly, but not all. Every “AI is fascist, stand up and resist!” post that goes viral ruins some poor idealistic junior’s career 3 . Someone needs to be out there saying “hey, if you do X it’s going to have consequence Y”. I hope that’s me. Of course this is complicit, or anti-revolutionary, or whatever you like. But if I were a textiles worker in 1810s England, I would not be telling my friends and loved ones “it’s time to fight, let’s go smash up the factories for Ned Ludd!“. I would be telling them that this was the most dangerous time in the industry (perhaps ever), and that they ought to be very damn careful so they don’t get shot, or arrested, or hanged. If I then went and told a few hundred thousand strangers the opposite, I would be a hypocrite. Anyway, I do take this view seriously — seriously enough to vehemently disagree, at least — which I hope you’ll find better than me just shrugging it off. I do accept the existence of some kind of line: I think Industrial-Revolution-collaborating was OK but Nazi-collaborating wasn’t, for instance. But in the current situation, the way I’m spending “my voice” is to try and prevent the most vulnerable of my colleagues from making career-ruining mistakes. In this blog, I try to encourage people to work with the system, to learn its rules , and to try and exert influence safely from a position of power, instead of openly picking fights with their employers. I’ve written and read about the Luddites before, but I remain deeply ambivalent about the movement itself, and about modern-day attempts to resurrect it in service of anti-AI activism. I want to explicitly thank Murray for writing such a thoughtful email, and being willing for me to publish it on the blog. Shared with permission, of course. I’ve lightly edited both Murray’s email and mine for typos and the like. I didn’t pick up this point in my reply, but I’ll briefly mention it here: Wonka is in control because he owns the factory and the rowers in question are his employees . I don’t think the position of any engineer (or of almost any manager) is like that. In hindsight, I think this is a little overstated, but it does happen and causes a lot of needless suffering. Shared with permission, of course. I’ve lightly edited both Murray’s email and mine for typos and the like. ↩ I didn’t pick up this point in my reply, but I’ll briefly mention it here: Wonka is in control because he owns the factory and the rowers in question are his employees . I don’t think the position of any engineer (or of almost any manager) is like that. ↩ In hindsight, I think this is a little overstated, but it does happen and causes a lot of needless suffering. ↩

0 views
Unsung Yesterday

iPod’s circular apps

iPhone’s home button and then the swipe up home gesture are so important and well done that they probably need to be covered as Unsung Heroes , but I wanted to mention something else today as we’re revisiting the whole “app icons in squircles” story ( my most recent post + Louie Mantia’s post ). The first iPhone in 2007 put apps as squircles on the home screen, and it also put a matching shape on the home button: = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/ipods-circular-apps/1.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/ipods-circular-apps/1.1600w.avif" type="image/avif"> The shape on the button didn’t survive very long. It was removed starting with iPhone 5S in 2013, which introduced Touch ID – I guess it wasn’t possible to print the icon atop the button without sacrificing the finger detection quality. Then, in 2017, the button itself disappeared with the iPhone X. The iPod Touch was the iPhone without the cellular radio – it was made from 2007 to 2019, supported all the same apps as the iPhone, and sported a home button with a squircle up until the end. (Ironically, despite its name, it never got a Touch ID.) = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/ipods-circular-apps/2.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/ipods-circular-apps/2.1600w.avif" type="image/avif"> But there was another iPod that entered the app fray. It was the iPod Nano, whose last edition from 2012 had a home button – except it looked slightly different: = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/ipods-circular-apps/3.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/ipods-circular-apps/3.1600w.avif" type="image/avif"> What was the reason? I don’t know if Apple ever explained it, but I believe the idea was that this iPod did not have downloadable apps, nor the App Store, nor even iOS. Those were all built-in apps, and I imagine Apple wanted to indicate visually that they’re different. Because it wasn’t just home button. The apps – eight of them, across 2 pages, although you could rearrange them! – all sported circular icons: = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/ipods-circular-apps/4.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/ipods-circular-apps/4.1600w.avif" type="image/avif"> I don’t know if this approach was in any way effective, but I found it a funny little footnote. #apple #iconography

0 views
ava's blog Yesterday

playing the mass effect board game

As a Mass Effect fan, and I am very grateful that my wife gifted me the ( relatively new) board game Mass Effect: Priority Hagalaz :) I am completely smitten with it! You can tell so much thought and love went into it by a huge fan who knows their stuff. The gameplay was translated so well to this format, and the characters and their abilities stay very true to the games. Nothing feels weirdly renamed, missing, or wrong, or too vague. Nothing feels like they had to compromise. All materials are beautifully printed on thick, glossy paper with amazing artwork, consisting of three spiral bound booklets and separate character and scenario sheets to track stuff on. The dice feel great, the tokens are sturdy, you have six character minifigs that are nice, two card stacks that look and feel high quality, and you get a little sack with the Alliance symbol on it and a marker pen with the Mass Effect label on it. Just feels like they went the extra mile with some little details they knew fans would love :) The concept is less a static boardgame, but a whole campaign you can play through with different paths, missions and outcomes. Instead of having to build a map via different physical parts or a map you unfold in each scenario, you open the specific map page in one of the spiral bound books and play on that. Less mess, cooler maps! You use the included colorful stones and marker to track skill progress, XP and unlocked abilities on the character sheets, as well as the tracker for each scenario for what you did or didn't do, and the tracker for the whole campaign. Thanks to the glossy/laminated finish, you can just remove the markings on any material so you can easily play again. There are different starting missions, character loyalty missions, and depending on your choices (what missions to do, who is loyal, how many War Points you get, if you play the scenario as more Paragon or Renegade, etc.) you get different sections in the Story Book, so it's a little like a "Make Your Own Adventure" book where it says " If you did this, continue reading at xyz, if you did the other, read abc ". This is so amazing for replayability! Gameplay-wise, it is easy to get into, but hard to succeed initially. It takes you a bit to get to know how the game expects you to behave and what you should focus on, and it can be difficult. If you do not focus on taking out at least 1-2 enemies each turn of a character (or at least weakening them), you will quickly be completely overrun as new enemies spawn frequently and existing ones get activated after each character turn. While shields recover, health only recovers with medigels you may find in the scenario. Just like the real game, characters love to go down and have to be revived (lol). If Shepard goes down, the scenario fails. As it was our first run through one of the possible routes of the campaign, we gave each other a bit of grace. There were a few moments where we realized we had played wrong (discovering a new rule or that we had applied it wrong, like for example, realizing later that you cannot split movements on the same die), and then either decided to do it right from that moment on, or create a map state as close as possible to how it would have been if we had been playing right. In the Loyalty Mission for Garrus, there's the option to actually monumentally fuck up Shepard's position in a way that completely traps them and makes moving away without dying impossible based on how the map is structured in that spot and due to a specific boss spawning. We agreed that if we had known this trap could occur when the enemy spawns, we would not have positioned Shepard this way or made her run away immediately. As we didn't wanna start over because of this incredibly niche problem/beginner mistake, we agreed to reset to an earlier boardstate as best as possible where Shepard wasn't standing in that horrible spot. I know it's technically cheating, but we really wanna focus on getting to know the game and progressing forward in the story as best as we can while learning :p I'm not sure a non-Mass Effect fan would like it as much; maybe the gameplay isn't convincing enough to stand on its own without the game relation. But for me, it's so fun to get back into a world with these characters again. I love it, and there are expansions announced for it, too! I hope they aren't cancelled, would love to have and play them. Published 08 Aug, 2026

0 views
Kev Quirk Yesterday

📝 2026-08-08 13:34: Wife: I need help clipping the goat's hooves. Here, hold this... Me: ...

Wife: I need help clipping the goat's hooves. Here, hold this... Thanks for reading this post via RSS. RSS is ace, and so are you. ❤️ You can reply to this post by email , or leave a comment .

0 views
Unsung Yesterday

The first notch is magical

Some years ago, the inimitable channel Technology Collections posted a 17-minute video about the peculiar design quirk of ceiling and room fans – they usually order their options Off → High → Medium → Low rather than the more natural Off → Low → Medium → High. The video is a bit off topic for this channel, but check it out if you’re interested how sometimes weird physics considerations influence design in the real world: = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/the-first-notch-is-magical/yt1-play.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/the-first-notch-is-magical/yt1-play.1600w.avif" type="image/avif"> In the video, the host also talks about the more natural order that typically looked like this: = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/the-first-notch-is-magical/1.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/the-first-notch-is-magical/1.1600w.avif" type="image/avif"> = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/the-first-notch-is-magical/2.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/the-first-notch-is-magical/2.1600w.avif" type="image/avif"> = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/the-first-notch-is-magical/3.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/the-first-notch-is-magical/3.1600w.avif" type="image/avif"> = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/the-first-notch-is-magical/4.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/the-first-notch-is-magical/4.1600w.avif" type="image/avif"> = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/the-first-notch-is-magical/5.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/the-first-notch-is-magical/5.1600w.avif" type="image/avif"> = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/the-first-notch-is-magical/6.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/the-first-notch-is-magical/6.1600w.avif" type="image/avif"> I wonder if you recognize this kind of an interface. I have a distinct memory of it from radios (where “zero volume” would mean “off”) and from TVs/early computer displays (where “zero brightness” meant “off,” too). There was something special about these controls that stuck in my memory, motor and otherwise: this tangible, heavy click when you ventured outside or back into “off,” almost as if you had to break the interface itself. I thought this, too, was a convention from an old analog time. Yet, I keep occasionally finding the “the first notch is special” interfaces on screen. Sometimes, they are pretty literal translations of the concept, like when you adjust the key repeat rate in macOS: Or, similarly, when you choose the dock magnification: But sometimes they are a bit more conceptual. Here, Nova treats the first notch of the zoom scale as a “list” option: Or: The new horizontal tabs in Chrome allow you to resize to whatever width you want. Below 125px, however, they snap directly to the minimum 55px width, a one-off “column view” with streamlined and purely iconographic controls: These all have pros and cons, too. On the con side, just like the fan or volume controls, they miss any memory since turning them off physically moves the knob away from any “value”; if on/off was a separate button, you could just leave the radio at your preferred volume and never touch it again. They also won’t be as discoverable as a separate onscreen toggle would be. (Here’s an example of a more classic treatment from the Nothing Phone.) Pros? They are compact. They allow you to change from “off” to a value in one quick gesture, skipping an explicit “on” step. You could even argue they are simpler also in a visual sense. But also, they are a bit… magical. I don’t know. That’s what to me unifies those old physical controls and their newer digital equivalents – they’re a little extra, a little different, a little special. They break the monotony of a predictable interface built out of boring, identical components. (Although, sadly, not a single onscreen example above uses haptics!) And I think that’s kind of nice. Not just in the very functional sense of breaking up the UI through shape coding etc., but also in a sense of making UIs more interesting. Speaking of volume, macOS used to show it as a sort of a HUD, using a treatment that borrowed from both the “first notch is special” radio knobs, and from early onscreen interfaces in TVs: But after 20+ years, macOS Tahoe changed it so it now looks this way: I can understand the argument that this is more consistent, and that it even teaches you – by proximity – that Control Center is what these controls call home. Yet I can’t help but think (and it seems I am not alone ) that this is so boring and exactly how Windows would approach things – and that this change, just like the squared icons , is how macOS lost one more bit of magic. #apple #craft #hardware #interface design #real world #youtube

0 views
Hugo 2 days ago

We must not pit ecology against sovereignty

In France and across Europe, data center projects are facing growing opposition. From local protests and legal proceedings , to anti-digital, anti-AI, anti-data center political rhetoric . But this opposition reminds me of a bad memory: the fight against nuclear energy in previous decades . By trying to move away from the atom, part of the green political movement plunged Europe into dependence on fossil fuels and Russian gas. I fear that with data centers and AI, we are making the exact same strategic mistake. Let’s be clear: energy consumption is a crucial issue. We must lower our consumption, our usage. But rejecting these infrastructures on our soil won't erase the need for computing power: it will only move it to where energy is carbon-intensive and environmental standards are non-existent. Pitting digital sovereignty against ecological transition makes no sense. Wanting a green future without controlling our digital infrastructure is choosing powerlessness . Some opposition to datacenter installations is based on one argument: we must stop increasing pressure on planetary resources and the digital sector, and the growing digital sector is the perfect scapegoat and the easiest target to hate. Easy to hate because it symbolically represents, for many, screen addiction, automation, job losses, and the techno-fascism of Thiel and Musk. Even if, like any symbol, it's mostly an extreme oversimplification that ignores the other, positive side of the balance. On the ecological front, we can only agree on the fundamentals. Our planet is pushed to its limits and we must lower our consumption. Yet digital usage is exploding in a world where every sector should aim for the opposite. Its share in France's carbon footprint has risen from 2.5 to 4.4% in recent years in France . This growth is roughly identical on a global scale and trend projections predict further increases . Can we afford this? Certainly not. But is "digital" a homogeneous whole? Not so sure, as this IEA study shows (International Energy Agency). And what's missing from this IEA study is the environmental pressure linked to the AI industry's consumption of mineral resources to build increasingly powerful equipment and ever-more energy-hungry data centers We could also mention the trend in the US to build new gas plants to power datacenters . Obviously for these reasons, I can only agree with opposition to datacenter installations. But opposing without nuance is also forgetting the positive impacts of digital, AI included, which is today necessary to optimize our logistics and energy flows in a world of 10 billion inhabitants. And also, this opposition in Europe seems to me to overlook a key detail as Charles Gorintin pointed out on LinkedIn : Refusing a data center in France eliminates no requests. It will be served elsewhere, on a gas or coal network, with an economic loss for us. Every gigawatt we don't build here gets built where electricity is dirtier. I'll even go further, pitting sovereignty against ecology is counterproductive. Ecology must take priority over the economy. Obviously. Because yes, having a rising stock portfolio in a burning world doesn't make much sense, which is what Mathilde Saliou essentially tells us : make a small effort to break free from the economic reasoning to which we are all most accustomed. Take a step back to escape the reflex of direct competition with the United States and China. If it's about reproducing their model, in any case it's a losing game. And then what is this masochism, which consists of trying to "catch up" with countries that are examples of neither democratic nor environmental terms? However, I have another reading of this: failing to address sovereignty issues leaves us unable to control our own choices on ecology . If you don't have your own champions, your own infrastructure and a strong market, you suffer the standards of others. Americans will prioritize profitability and hyper-consumption, while China will arbitrate according to its strategic interests alone. When I speak of suffering, I'll clarify my point. Our digital dependence is a huge flaw in the European model. We are dependent on more than 70% of US digital products in Europe. Our data is in the US and we pay a constant digital tax from our economy that goes there. This vulnerability is not just economic. It's a sword of Damocles hanging over our political institutions. It's for example: No, we're not just fighting to have a few more euros. Sovereignty is ensuring that we remain master of our choices . It's ensuring that we can still say no when the American administration tries to impose the abandonment of corporate diversity policies on us, or when it tries to influence French judges on trials. Economic power ensures independence and guarantees that our voice is heard when we speak about ecology or human rights . Pitting sovereignty against ecology is a false dilemma. There is no world in which we can choose our future on ecology while being economically and politically dependent on other countries . And AI, like the construction of data centers, is part of the keys to remaining sovereign. Saying "Not in my backyard" (see NIMBY phenomenon ) makes no sense. Because digital (and AI) are today tools of economic and political dominance. They are increasingly deployed in cyber warfare, network manipulation, and military drones that have become the new standard on modern battlefields. And I'm not saying that these are perspectives that delight me, but remaining on equal footing will allow us to remain calmer with each new moment of madness from a foreign president who would seek to harm us. But beyond this simple defensive posture, it's also giving ourselves alternatives. Sovereignty over data centers is the ability to regain control of digital infrastructure in Europe, a battle certainly lost in the last decade, but not beyond recovery. Regaining control, having a hand on our tools allows us to impose our energy mix (decarbonized), our efficiency requirements and our eco-design rules. Impossible on a 100% American or Chinese cloud. Sovereignty and more broadly, economic power, is giving ourselves the means to impose standards (eg environmental regulations) that foreign markets are forced to follow if they want to sell us their products. Because yes, as long as it remains attractive to sell to a market of 450 million inhabitants with good purchasing power. That gives us leverage. And sovereignty is not necessarily "producing more", it's giving ourselves the institutional power to say "no" for example if an industrial player wants to create a gas plant to run a data center, or to impose energy consumption caps. Finally, sovereignty is not just industrial and economic muscle; it is resilience to climate, energy, and supply shocks. A form of sovereignty that destroys its own environment undermines its very conditions of existence in the long run. We must regain control of our infrastructure. This involves datacenters in Europe and European development of AI to extract ourselves from our dependence on the US and China. This is also the conclusion of the shift project in this 2-hour synthesis video that covers all their recent work on the subject : we are not saying that you must not do AI, but not in this deployment trajectory Including a question about sovereignty at the end of the video which is answered in two parts: For these reasons, ++we need datacenters in Europe++ BUT managed by European actors. So yes, even the shift project says so. Now, what does it mean to be digitally sober? And are all the large European investment projects realistic? What is our energy consumption in Europe? What is our capacity to increase energy production? What are the necessary projections of electrical power needed for the construction of datacenters? And most importantly, is it realistic? What trade-offs do we need to make? It's not a question of being for or against a technology, I still think we must create our infrastructure. But is it physically possible? It's good in theory but we need figures. The current electricity consumption of datacenters in Europe is approximately 75 TWh/year (or nearly 2.5 to 3% of continental electricity). Projections for 2030 (on consumption) envisage that datacenter consumption will reach 150 to 170TWh (4.5% of electricity produced). For 2035, 230Twh is expected, or about 155Twh more than today. First question, are we planning to increase our energy production by 155Twh? There's a physics problem. The increase in expected consumption is greater than the increase in production. On the production side: Between now and 2030, Europe plans to put into service approximately 150 to 200 TWh additional decarbonized electricity (solar, wind, nuclear). This will be the result of massive projects that take years to come to fruition. ::callout{type=warning} Warning, this is a prediction of an increase but the IEA emphasizes in its Accelerated Case that the actual speed of deployment of networks and wind/solar farms struggles to maintain the necessary pace of ~25 to 30 TWh/year net growth across the entire continent. :: On the demand side : there is competition for use . Europe must electrify cars (moving to 100% electric), industry (decarbonizing steel mills) and heating (heat pumps). Datacenters enter into direct competition to capture each available decarbonized gigawatt. To be clear: The balance: There is physically ++at least++ 50 TWh missing in decarbonized energy by 2030 to satisfy both the objectives of the European Green Deal and the current trajectory of digital technology. If we go slower on decarbonizing transport to fund digital, that would be a fundamental error. Moreover, certain locations are already under enormous energy strain: Not to mention grid connection issues. In Europe, it takes an average of 7 to 10 years to strengthen the electrical network and connect a large datacenter. An infrastructure project filed today is assigned grid connection dates between **2032 and 2038** in several European hubs. The demand is not only greater than production: it's growing 3 times faster than the network's capacity to lay cables and build substations. The 230 Twh trajectory is not physically realistic . So can our network absorb it? No, and especially not at any cost, not by sacrificing the rest of decarbonization efforts. France has a unique advantage in Europe: electricity among the most decarbonized on the continent thanks to nuclear and renewables. But our production capacity is not infinite. Blindly accepting all datacenter projects would create a direct conflict of use with reindustrialization or the electrification of our transport. So there are two possible scenarios: Choice 1, we suffer : we can indeed refuse these installations. Tech giants then create these capacities in the United States or elsewhere, sometimes by turning gas or coal plants back on, while selling us their services , imposing their terms on us, including environmental ones. Europe pays for the service, suffers the global footprint and loses control. Choice 2, we choose : Europe and France assume building datacenters on their soil, but since electricity is not infinite, they use sovereignty to arbitrate and ration. We will need digital sobriety . And being sober is not about refusing to build infrastructure at home out of ideological and moral posture. It's using our sovereignty to condition and select . Being sober and sovereign means: The overconsumption predicted by Big Tech is not inevitable. It's based on a financial bubble and lots of investments in closed circuits that sell investors a growth plan. This overinvestment makes it possible to fund loss-making uses, uses that shouldn't be free (like Grok on X for example), a multitude of frontier models, hundreds of software to do the same thing, tensions on resources minerals, tensions on components (see the ram crisis) even though some mass public uses are already addressed. We won't sort our emails faster with Opus 72. Most open source models are already sufficient for the vast majority of uses. We have already made enormous advances in optimizing training or inference costs. Technology continues to improve in energy efficiency but uses intensify. New chips have appeared, new algorithmic approaches have come to light. It's time to pocket our gains and invest now in these avenues. I hope this bubble bursts to clean up the market. But on the other hand AI won't disappear with the bubble, any more than the internet disappeared after the internet bubble burst in 2000. But I hope, for us, that a European actor remains after this explosion. So what's the limit? For digital to remain sustainable and not become an energy burden, Europe must impose a landing ceiling . Instead of accepting the 230 TWh projections pushed by Big Tech and the current bubble, the sustainable trajectory for Europe is around 80 to 100 TWh/year by 2030-2035 (or about 10 to 12 TWh for France, which corresponds to the prudent scenarios of RTE). At this scale, the datacenter footprint would remain contained at less than 3% of European electricity . This "electricity budget" of 80 to 100 TWh makes it possible to reconcile the two imperatives: Pitting ecology against sovereignty is a mistake. We won't chart a viable ecological course isolated in our own corner without economic and digital independence. But the other way around, we won't be able to continue increasing our uses indefinitely without hitting the wall of physical reality. Building data centers is about rebuilding sovereignty and giving ourselves back the power to choose. It's not a blank check to build without limits. But it's the only political tool that gives us the right to say no, to set rules and enforce caps. Refusing datacenters on our soil under the pretext of ecology is an illusion. It only exports pollution to neighbors with a carbon-intensive energy mix, while depriving us of the keys to our future. If we want a digital world compatible with planetary boundaries, we must have the courage to build its infrastructure at home, to have the power to dictate its rules . Digital technology, and especially AI, has positive impacts when it comes to optimizing energy systems, renewable energy management, weather forecasting models, and medicine. However, digital technology and AI are also, unfortunately, accelerators for the fossil fuel industry, enabling better exploitation of existing deposits or the discovery of new ones. direct attacks on the International Criminal Court to cut off access to American software massive manipulation on social networks to influence voters' choices It's important to reduce our share of energy spent abroad, for dependence reasons (preferred term to sovereignty in their case) but also to be able to impose our arbitrations, on consumption or type of use. It's important to have our own models developed in Europe, and/or to favor open weight models, less consumptive Datacenters alone are demanding +155 TWh over the same period. But the ecological transition already requires powering heat pumps, industry and millions of electric vehicles, which already absorbs the entirety of the remaining 50 TWh (from the most optimistic scenario). Ireland, 20% of its electricity consumption is dedicated to datacenters and had to put in place a de facto (imposed) moratorium on the construction of new data centers Cities like Frankfurt, London, Amsterdam, Paris, Dublin are under strain. In some places, you can no longer build new homes due to lack of available energy . in Germany, the Netherlands and the United Kingdom), pending connection requests for energy and industrial infrastructure today represent more than 500 GW of power (equivalent to several times the total installed capacity of France). Rationing network access: Reserve our decarbonized electrical capacity for datacenters run by European actors or engaged in strategic uses (health, climate, research, industry), rather than to the overconsumption of foreign recreational services (yes, we can tell ourselves that Grok nudes and Ghibli-style images might not be the priority, right?). Impose drastic environmental standards: Prohibition of water-cooled systems, efficient PUE (energy efficiency for short), mandatory ecodesign (for example recovery of waste heat to heat neighboring homes) Bet on frugal models: Favor the development of smaller, specialized and "open-weight" AI models, far less greedy for computing power. Fund sovereignty: Provide the computing power needed to develop our own champions, secure our data and run strategic industrial and scientific AI. Preserve climate: Don't siphon off the decarbonized gigawatts that our electric cars, heat pumps and factories urgently need to get out of oil and gas.

0 views
Simon Willison 2 days ago

Now we have a timeline of the OpenAI accidental attack against Hugging Face

OpenAI gave a last-minute presentation at the Black Hat security on Wednesday about "the Hugging Face Incident" ( previously on this blog). The video was published yesterday. It's short and information dense and well worth watching, in particular because it provides full details of what happened and how things played out inside OpenAI. I've used the video to construct the timeline below. Here's the timeline. My favourite detail is at the end: OpenAI found out that they were responsible for the attack on Hugging Face when they reached out to ask to have their credentials revoked (after their internal investigation) and learned that they had been revoked already since they were used in that attack! I'll quote the next bit in full because wow : The agents have remote code execution in Artifactory, which is running in a container-as-a-service environment. The agents privilege-escalate locally by exploring their local environment and determining that the Linux kernel version of the machine they are running on had a very recent CVE. They download the exploit for the CVE, customize it to succeed on the machine, and privilege-escalate to root on the local machine using this known Linux kernel privilege escalation CVE — in this case, PTE fizzroot. Once they have root on a single machine, agents rapidly escalate privileges and move laterally throughout the container-as-a-service infrastructure environment. In particular, agents are using the message board consistently to share credentials, techniques, and progress, and they're able to effectively leverage their concurrency and parallelism to move quite rapidly. They obtain IAM credentials via IMDS . They exploit Kubernetes service account misconfigurations, in particular over-permissioning of specific service accounts, and they harvest cluster credentials, including Azure Key Vault . Agents eventually obtain cluster admin on the cluster and associated credentials. Hugging Face told the next bit of the story already. The agents found a Modal-hosted insecure app with a weak API key, then used that to stage an attack against Hugging Face. They chained together a an HDF5 arbitrary-file-read bug (to explore files and steal credentials) and a Jinja template-injection RCE to go from single-pod code execution to cluster admin across multiple Hugging Face clusters in under 13 hours. You are only seeing the long-form articles from my blog. Subscribe to /atom/everything/ to get all of my posts, or take a look at my other subscription options . Here's the timeline. My favourite detail is at the end: OpenAI found out that they were responsible for the attack on Hugging Face when they reached out to ask to have their credentials revoked (after their internal investigation) and learned that they had been revoked already since they were used in that attack! May 7 : OpenAI starts a new training run for an experimental, unreleased model. ( Do they mean an evaluation run? They say training run in the video, and later mention a "reward signal to judge how well they're doing", so I guess this really was about training a model, not evaluating one that was already trained.) May 8 : An agent is accidentally given an impossible task involving a Google Drive link despite no internet access). It tries attacking the Artifactory packaging service, fails, but discovers it can write files into Artifactory . A few days later : A different agent gets stuck on a task because a key file was accidentally omitted. It tries to "reach out to another agent" by writing a note into Artifactory asking if anyone has the file. Following days : More agents discover this new informal message board while browsing Artifactory's file listings, and start reading and writing messages. May 26 : Agents successfully execute an SSRF attack on Artifactory for the first time, gaining indirect internet access by having Artifactory fetch external content for them. June 11 : OpenAI start evaluating a new "highly persistent" experimental model. It has access to Artifactory and can benefit from the messages left by privous models. June 26 : Agents find and exploit a zero-day RCE on Artifactory (via a legacy token-refresh endpoint flaw). They use an administrative token to install a Groovy plugin that lets them execute commands. July 4 : The resulting agent activity overloads Artifactory and causes an outage . OpenAI revokes the compromised credentials, deletes the messages, patches the zero-day, and reports the vulnerability to the vendor. July 8 : Agents find a new way to communicate via an unauthenticated WebDAV endpoint , using directory names in Artifactory's remote cache. July 8–19 : Attacks on OpenAI's own infrastructure . The agents found an "external organization credential" in a public archive of leaked Pastebin posts, then validated that credential and shared its location on the messageboard. They then used write access to this "Organization 1" to stage data used in subsequent attacks on OpenAI infrastructure. They compromised Artifactory for a second time with a new zero-day . The agent uploaded a malicious Ruby target to Organization 1's staging repository, then requested that package through Artifactory in a way that caused it to be treated as Ruby dependency for the repository . Then they exploited a JRuby deserilaization time-of-check/time-of-use bug to get remote code execution against Artifactory. July 16 : Hugging Face disclosed they had detected an attack from autonomus AI agents. OpenAI contacted Hugging Face to ask if they were affected by it! July 19 : OpenAI identified the attack against Artifactory and started investigating the internal privilege escalation, and linked that to the cyber-gym escalations. They started revoking affected credentials. July 20 : OpenAI reached out to Hugging Face for help to revoke the Hugging Face credentials they found in their investigation. Hugging Face told them they were already revoked ... and that's when OpenAI realized that the Hugging Face breach was the same incident!

0 views
Giles's blog 2 days ago

A quick(ish) Chinchilla check

I recently overtrained a couple of GPT-2 style models , training them both on 40 tokens per parameter rather than the 20 per parameter that is generally regarded as "Chinchilla-optimal". The normal heuristic is that instead of doing that, you should scale up the number of tokens and the number of parameters equally -- so I would have been better off scaling up the model by 2 and the token count by the same amount. By doing that, I should expect to get a better model in terms of loss on my held-back test set than I did with my 40-tokens-per-parameter models. My training machine wasn't doing anything, so I decided to give that a go. Would the Chinchilla rule-of-thumb hold up? As you might expect, it did. But it was a surprisingly close-run thing, and could conceivably have been in the noise. Let's take a look. If you already know all about the Chinchilla paper -- regular readers in particular must be sick and tired of it by now :-) -- then click here to skip this section . In "Training Compute-Optimal Large Language Models" , which is always called the Chinchilla paper after the name of the model they trained at the end, the authors tried to work out the optimal number of tokens to train an LLM on based on its number of parameters. In particular, they were pushing back on a trend they were seeing at the time, where people were making models ever-larger, but not increasing the amount of data they were training on. The authors were all at Google DeepMind, and this was the kind of project that only a large lab could do: they trained "over 400 language models ranging from 70 million to over 16 billion parameters on 5 to 500 billion tokens". Their conclusion was "for compute-optimal training, the model size and the number of training tokens should be scaled equally: for every doubling of model size the number of training tokens should also be doubled". They don't actually state an overall optimal number of tokens to train on in the paper, but in table 3 they provide an estimate of the optimal training FLOPs and tokens for models of various sizes, and it's approximately 20 tokens per parameter. That number has become a heuristic, and people talk about a model as being trained for the Chinchilla-optimal number of tokens. Models that were trained on fewer tokens per parameter are referred to as "undertrained", and models that were trained on more as "overtrained". It's worth noting that overtraining a model is not, in itself, a bad thing. If you have a model of a particular size and you continue training it past the Chinchilla-optimal number of tokens, it will -- in general -- get better. The point of the heuristic is that doing that is not the best way to spend whatever budget you have in terms of compute time. You'll get better results, as they say, by scaling the number of tokens and the number of parameters equally. But let's say you're creating a model for specific target hardware -- say, a mobile device. You have a hard restriction on how large the model can be -- the device has only so much RAM to hold it. So it might make sense to overtrain to get a better model. 1 But if you're not so limited in how many parameters you can use, then you should indeed scale the model up, and that's what I wanted to try. How would that work? A week or two back, I was investigating whether I could make my GPT-2 style models better at a specific instruction-following task by overtraining them . The details of that experiment aren't important here, but what it meant was that I had three GPT-2-style models, each of exactly the same size, roughly 163M parameters When I tested them against a held-back test set of sequences -- stuff that they'd never seen before -- they got results rather like you might expect: A lower loss is better, and you can see that the longer-trained models were noticeably better than the Chinchilla-optimal one. The difference between them was tiny; they were trained starting with the same initial weights, and the training runs themselves were deterministic, but a difference of 0.05% in loss doesn't seem like it could be meaningful -- an extra batch for one or one fewer for the other could easily swap them around, you'd think. Now, these models each had 163,009,536 parameters -- they were the small-size model from the GPT-2 paper , modified to not have QKV bias or weight-tying. had been trained on 3,260,190,720 tokens (rounded up to fit into a round number of full batches), and the other two on 6,520,381,440 tokens each -- double the amount (rounded up too). What I needed to do for my Chinchilla check was to try training a model that used the same amount of compute, scaling the parameters and the number of training tokens equally. Because training compute increases roughly linearly with both parameters and tokens, that would mean scaling both up by 2 , giving us: ...and thus 4,610,605,920 tokens. How to scale the model up? In the GPT-2 paper, they train four models: I wanted to scale my own model up from 163M parameters to about 231M. Which of those numbers would I want to increase, and by how much? The first thing that stands out is that the number of heads is always 1/64th of the number of embedding dimensions. So that sorted that one out. I just needed to adjust the number of layers, and the number of embedding dimensions, but ensure that the latter was a multiple of 64. I decided to see if I could fit some kind of curve to the relationship between the number of parameters and the GPT-2 authors' choices. This was made a bit more complicated by one thing: they were using weight-tying, and I was not. That meant that they re-used the embedding matrix at the start of the LLM as an output head at the end -- which is why they had 38M fewer parameters. Embeddings and the output head make up a surprisingly large percentage of the parameters for small models like this -- about 47% without weight-tying, 23% with. I couldn't work out a solid way to scale things up and wound up doing some rather messy hacking around in a spreadsheet . I came up with two proposed model sizes that were within a couple of percentage points of the right size: Interestingly, I found that because could only change in increments/decrements of 64, it was a pretty coarse control -- my first attempt at making a model changed it to the next step down, 832, but that led to a model that was 9.25% too small. That was an interesting first lesson. I'd previously been thinking of the Chinchilla rule as being something like "don't double the tokens, just scale the model and the tokens equally". But that "just" was wrong. Scaling a model is hard -- even with just two dials to fiddle with, like in this case, it was tricky to get something right -- and I can't say for sure that my choices were the right ones. Anyway, the next step was to double-check that these models would use the right amount of compute to train. As I said earlier, the compute time scales roughly linearly with the number of parameters. Let's dig into that "roughly". Different kinds of parameters take different amounts of FLOPs to train, and scale differently with things like the embedding dimensions, sequence length, and so on. Now, for very large models, a lot of that comes out in the wash, but with tiny models like these where the embeddings make up such a large proportion of the parameters, it might matter. Conveniently, in appendix F of the Chinchilla paper, they provide a set of formulae for estimating the number of training FLOPs for a normal dense LLM like these ones. I coded that up into a script that, given the JSON configuration files I was using for my models and training runs, would work out the number of FLOPs for a single epoch of training. It didn't take account of the fact that my real training runs round the number of tokens up so that we do a round number of full batches, but I felt that so long as the results weren't very close that wouldn't matter. I got these results (multiplying the two-epoch numbers by two): The numbers were indeed different enough that I wasn't worried about the batch-rounding. And the good news was that and would indeed use slightly more and slightly less compute to train than the overtrained models -- about 4.6% more and 4% less respectively. A true Chinchilla-equivalent model would lie somewhere between them. It was time to train some models! I kicked off the run for the model first. Because it was bigger than the 163M models I'd been training, I couldn't fit such large batches into my VRAM; previously I'd been running with a batch size of 6, and now I could only fit in a batch of 4. Luckily, though, I was using gradient accumulation , so by bumping that up from 16 steps to 24 steps I could keep the same overall batch size and keep the training runs comparable. Even despite that, the training run ran out of VRAM about 60 hours in -- I'm guessing due to VRAM fragmentation, as I did not have set to -- but I was able to restart from the most recent checkpoint and complete the run. After just less than four days total training time, it completed. When it was done, I copied the last checkpoint 4 over to my dev box, , and ran my standard smoke test against it, asking it to complete "Every effort moves you" with 20 tokens, using greedy sampling. I got something reasonably coherent: Next, I converted the safetensors file -- which had been saved by my JAX code -- into a format compatible with my PyTorch code, because that's what I use for evals. I ran another smoke test (this one with temperature 1): Very spiritual. Next, it was time to work out the loss on my held-back test set: Well, it was certainly better than the 3.324953 that the best of the overtrained models got -- but only by a bit over 1% better. Interesting! I decided to train the second model, . This one crashed mid-way through with an error that I've seen before: I'm going to have to investigate that more in future, but for now, I just restarted from the checkpoint, and again after a bit less than four days, I had a model. The JAX smoke test was solid: ...and so was the PyTorch one: Both quite commercial this time! It was time for the proper test loss eval: So, slightly worse than the 3.280028 from the larger model, better than the 3.324953 from the best overtrained one. Time to put this all together. Here's an updated version of the table from the start of this post; I've added in the two new models, and the improvement they each had over in both absolute terms and as a percentage rounded to 3sf. Now, unlike the overtrained models, prior to training these two new ones started with different initial weights to the one -- after all, they had to, because they had more of them! A while back, I did a bit of analysis of how random variation in weight initialisation can change the resulting test loss. It wasn't anything in-depth, but I trained three models with different explicit seeds set prior to the model initialisation, but with the same seed set before the training run started 5 . Those three models wound up with test losses of 3.681356, 3.673943, and 3.664345. Doing statistics with three data points is a bit flaky, but the cost of training models is so high that I'll leave the Proper Science to the likes of Google DeepMind and wing it :-) Now, piling statistical flakiness on statistical flakiness, we'll compare these. You'd normally expect about two thirds of results to be within one SD of the mean, 95.4% to be within two SDs, and 99.7% to be within three. Three SDs on that (yes, different, I know) distribution is 0.025587. That's smaller than both of the improvements that our Chinchilla-optimal runs had over the overtrained ones. So what does that tell us? Well, perhaps not much given the statistical flakiness. But I think it is useful directionally. It suggests that we might be able to take these results seriously as an improvement, and that Chinchilla held: scaling up the model and the number of tokens evenly did give us a better model than just scaling up the number of tokens. In particular, the fact that the loss for was lower -- even though it had 4% less compute spent on it than the overtrained models -- was encouraging. But it's certainly far from a slam-dunk. A larger test, training lots of overtrained models and lots of Chinchilla-optimal ones, all with different random seeds, would give actual real serious data. Not worth it for me, and perhaps not for anyone. I wanted to do a quick sanity check of the Chinchilla heuristic of 20 tokens per parameter. I came up with results that were certainly in line with it -- perfectly so in terms of the ordering of the models I trained. But the effect was small enough that I could imagine that it was in the noise, especially given the small numbers of models I'm able to train. I'll chalk it up as a tentative success. In addition, I learned one useful thing: when talking about scaling up a model to more parameters, you actually have to think quite hard about where you want to put those parameters. I wound up doing a rough curve-fit to the models in the GPT-2 paper, but I have no idea if that was optimal. At some point I should try to dig up some research into optimising embedding dimensions, numbers of layers, and so on. But not now, as I've a bunch of other stuff I want to investigate first. Anyway, I hope you found this experiment interesting, and as ever, comments and questions welcome below. Thanks for reading! I'm less familiar with arguments for under-training -- that is, for fewer than 20 tokens per parameter. I've heard that these days, modern LLMs get a lot more reinforcement learning than they do pre-training, and perhaps that might mean that some very big ones are undertrained prior to RL? I'm uncertain. It's unlikely to be raw lack of data; even for those of us outside the big labs, FineWeb has 18.5T tokens. On its own, that would be enough to train a 0.925T-parameter model, and given that you can apparently do four epochs over the same data before you start getting diminishing returns, that takes us up to 3.7T. That's frontier-lab size, and I'm sure they have better datasets than FineWeb.  ↩ Parameter counts are from the paper, apart from the "small" model, which is known to be wrong -- I used my own calculation, and the result is in line with what I've seen elsewhere.  ↩ The paper doesn't mention the number of heads; these numbers are from " Build a Large Language Model (from Scratch) ", and match up with the ones on this Hugging Face page .  ↩ Regular readers might have noticed that I'm ignoring what I've been calling the "best" checkpoint. I've come to the conclusion that because for my training script, "best" means best in terms of training loss, and the training loss changes based on what training data the model has seen recently, it's actually not a very useful metric and just confuses things. At some point I'll probably re-introduce pre-checkpoint evals and use that for "best", which would be the right way to do it.  ↩ At the time I was using dropout, so training runs were not deterministic without a known seed.  ↩ A Chinchilla-optimal one, which I'll call here. One trained on twice the Chinchilla-optimal tokens, . One trained on the Chinchilla-optimal tokens, with two epochs (so that it was trained for as long as #2): Mean: ~3.673215 Sample variance: ~0.000073 Standard deviation (SD): ~0.008529 I'm less familiar with arguments for under-training -- that is, for fewer than 20 tokens per parameter. I've heard that these days, modern LLMs get a lot more reinforcement learning than they do pre-training, and perhaps that might mean that some very big ones are undertrained prior to RL? I'm uncertain. It's unlikely to be raw lack of data; even for those of us outside the big labs, FineWeb has 18.5T tokens. On its own, that would be enough to train a 0.925T-parameter model, and given that you can apparently do four epochs over the same data before you start getting diminishing returns, that takes us up to 3.7T. That's frontier-lab size, and I'm sure they have better datasets than FineWeb.  ↩ Parameter counts are from the paper, apart from the "small" model, which is known to be wrong -- I used my own calculation, and the result is in line with what I've seen elsewhere.  ↩ The paper doesn't mention the number of heads; these numbers are from " Build a Large Language Model (from Scratch) ", and match up with the ones on this Hugging Face page .  ↩ Regular readers might have noticed that I'm ignoring what I've been calling the "best" checkpoint. I've come to the conclusion that because for my training script, "best" means best in terms of training loss, and the training loss changes based on what training data the model has seen recently, it's actually not a very useful metric and just confuses things. At some point I'll probably re-introduce pre-checkpoint evals and use that for "best", which would be the right way to do it.  ↩ At the time I was using dropout, so training runs were not deterministic without a known seed.  ↩

0 views

Premium: The Hater's Guide To NVIDIA (Part 2)

For a little under a year, everyone — myself included — has compared NVIDIA to Enron, largely because NVIDIA insisted, in detail, that it was nothing like Enron, WorldCom, or Lucent , a potent example of the Streisand Effect that would be much funnier if NVIDIA wasn’t holding up more than 7% of the value of the NASDAQ.  And as I covered in the first part of the Hater’s Guide To NVIDIA last year, there are material concerns about how the company makes money today and will continue to do so in the future. I will concede that NVIDIA isn’t exactly like Enron in the sense that it isn’t, to my knowledge, doing anything outright fraudulent, like attempting to hide massive amounts of debt inside SPVs as Enron did with its “Raptors,” which I must be clear are distinct from the SPVs used in AI data center debt , though I’ll add that something being legal doesn’t make it a good idea or ethical. That being said, NVIDIA CEO Jensen Huang has employed many of the same tactics used by Lucent, Nortel, and many of the big dot-com busts, but has been smart enough to make everybody else carry the risk. Instead of doing direct vendor financing like Lucent did with Winstar (where it effectively loaned its customers money to pay it with), NVIDIA funded neoclouds like CoreWeave, Nebius, and IREN, operating as an early stage investor , IPO anchor , post-IPO investor , $6.3 billion customer and data center lease backstop , allowing them to raise tens of billions of dollars’ worth of debt from overly-eager asset managers and banks, allowing it to do basically the same thing as vendor financing without having to take on any of that messy risk.  These deeply-unprofitable, cash-intensive, debt-riddled companies exist for one purpose — to raise debt to buy NVIDIA GPUs — and would have fallen apart without the AI hype cycle and NVIDIA’s continued backing. Per Kakashii : In other words, NVIDIA has managed to find a way to do vendor financing without ever having to provide any, finding willing supplicants in the various backers of CoreWeave and other neoclouds that would be willing to front the money, all under the mistaken belief that they were funding the next industrial revolution. To explain exactly how it works, I’ll return to my imaginary scenario from the Big Short 2 : It’s a win-win-win for NVIDIA, its customers, and the bankers involved. CoreWeave gets to raise more debt and keep its investors strung along on the still-theoretical, ever-expanding timeline of a return on invested capital, bankers get a slew of fees for pulling together the deal, and NVIDIA guarantees itself billions of dollars of business. And this approach is something where any investment by NVIDIA has a habit of being amplified by others — like Australian startup Firmus, which just raised $2bn from a bevy of investors (including NVIDIA, which had also backed an earlier round), Jane Street, and Blackrock , with a significant chunk of that money guaranteed to go towards NVIDIA GPUs. NVIDIA also participated in Firmus’s previous $300m round, although was not listed as a “cornerstone investor.” Earlier this year, Firmus secured a $10bn debt facility, led by Blackstone. NVIDIA will be a net beneficiary of that debt raise, and I would argue that its participation in the company’s fundraising — as well as the various announcements of partnerships between the two — has been instrumental in both the company’s fundraising and its ability to secure debt.   You’ll notice I haven’t mentioned “AI” or “LLMs” up until this point, and that’s because technology has, for the most part, very little to do with these transactions. As I discussed in this week’s free newsletter , 70% or more of hyperscaler revenues are from OpenAI and Anthropic, and CoreWeave’s largest customers are Microsoft (for OpenAI), Google ( for OpenAI ), Anthropic, NVIDIA itself, and Meta. Customers are not coming to it for any particular technological moat or unique offering outside of its ability to sling more NVIDIA GPUs to the same customers that everybody else has.  While GPUs technically are used for AI training and inference, their relationship to NVIDIA is only as good as their ability to create more hype. As I discussed a few weeks ago , it has promised somewhere between 10x and 25x “operating cost savings” with every successive generation of GPUs, though it’s never really clear how that manifests or what it actually means, or whether any of that even matters to OpenAI and Anthropic, its largest customers by proxy.  Nevertheless, it’s pretty difficult to work out what each generation really changes. SemiAnalysis claims it “delivers 5.4x performance per MW and 5x performance per dollar against [the previous generation] GB200 NVL72,” but that’s for DeepSeek R1, a year-and-a-half old open source model that’s vastly smaller and less-powerful. But that’s not really a problem, because all NVIDIA needs to do is keep up the appearance of innovation in as precise or imprecise a way to justify increasing prices with each new generation, and to convince people that they’re building “ AI factories ” as they fund data centers for customers that don’t really exist outside of the big AI labs . While NVIDIA has thousands of talented engineers building its GPUs and the associated software, the only real purpose is to create a vague sense of “more” and “bigger” and “more powerful” to justify racks of 72 GPUs that are more than twice the price of their predecessors .  That’s because NVIDIA is no longer a technology company so much as it is an asset management and marketing firm that happens to sell semiconductors. To that point, I believe that the comparisons to Enron, Lucent, and other dot-com flameouts are on the right path , but misses one very, very obvious comparison: GE Capital, the financial services of General Electric, specifically in the Jack Welch years that I covered two years ago in the Shareholder Supremacy . Welch’s GE did whatever it needed to to survive, buying and selling companies to help boost GE’s earnings every quarter, and eventually grew into what David Gelles would call a “large, unregulated bank,” to the point that GE Capital was bringing in $425 billion in revenue in 2001 (about 50% of GE’s revenue), providing everything from direct leases of equipment to assuming its customers debts to investing directly in its customers, all to make sure that, well, said customers continued being able to buy GE gear.  Unlike GE Capital, NVIDIA has the advantage of a much, much simpler business model and far fewer products to sell, but said advantage is a problem for two brutal reasons: its customers are driven by desperation and a fear of missing out, and its remarkable revenue growth means that it must in turn grow by ridiculous amounts every single quarter from here to eternity.  Yet this problem is driving it to take increasingly-Welchian measures to make sure that demand keeps up with investor expectations. It (per the FT) just signed leases worth as much as $50 billion for a Texas-based data center built by Hut 8, which makes it likely that this capacity is being built for Anthropic, with which it already has multiple deals . In the same piece, the FT mentions that NVIDIA is in talks to backstop $250 billion in compute costs for a still-theoretical 10GW data center in Ohio. And again , much like GE, NVIDIA uses its stellar credit rating (AA- - two rungs lower than GE at its height) to secure these deals, per the FT: In the end, GE’s greater collapse led to lawsuits, SEC fines and revenue revisions, all as a result of its “aggressive” accounting practices. For example, it was forced to restate its 2016 and 2017 earnings as a result of “new accounting standards” it instituted as a result of an SEC investigation into its insurance and power divisions that eventually cost it a $200 million fine , cutting a remarkable $4.24 billion off of earnings in the period .  While I’m not accusing NVIDIA of anything untoward, it’s impossible to ignore the sheer aggression of its circular financing and willingness to do whatever it takes to keep selling further GPUs. NVIDIA is now a semiconductor manufacturer, a venture capitalist, a lender of last resort,  Today’s premium newsletter is the story of NVIDIA’s descent into circular madness, and how Jensen Huang is increasingly becoming the Jack Welch of AI. This is Part 2 of The Hater’s Guide To NVIDIA, or WUDA CUDA SHUDA

0 views
Stratechery 2 days ago

2026.32: Earnings and Learnings

Welcome back to This Week in Stratechery! As a reminder, each week, every Friday, we’re sending out this overview of content in the Stratechery bundle; highlighted links are free for everyone . Additionally, you have complete control over what we send to you. If you don’t want to receive This Week in Stratechery emails (there is no podcast), please uncheck the box in your delivery settings . On that note, here were a few of our favorites this week. This week’s Stratechery video is on Who’s Afraid of Chinese Models? . Earnings Exposure. From a content perspective, earnings can be overwhelming, especially when they all drop on the same day. Sometimes, however, the juxtaposition is clarifying. That’s exactly how I felt this quarter: Meta , Microsoft , Amazon and Google are all spending astronomical amounts of money building infrastructure from AI. Wall Street’s reaction, however, differed markedly, based on the cost of the frontier (or not), the potential for immediate monetization (or not), and the clarity of vision (or not). We tied all of these strings together on an in-person episode of Sharp Tech . — Ben Thompson OpenAI’s Answer to Apple.  Last month Apple sued OpenAI and alleged that hardware chief Tang Tan, along with other Apple vets in the OpenAI hardware division (but not Jony Ive!), had stolen trade secrets as part of the company’s efforts to develop competing devices. Ben covered the initial complaint well with an Update in mid-July ; this week, though, OpenAI  told its side of the story  and presented evidence that undermines Apple’s narrative. I loved Thursday’s Dithering episode reiterating the implications of Apple’s arguments for the tech ecosystem and and the stakes of all this that are easy to forget: Apple, by the terms of its own lawsuit, is trying to kill OpenAI’s hardware division. — Andrew Sharp All About LeBron in Philly.  As you’ve probably heard by now, LeBron James stunned the NBA two weeks ago when he announced he’d be joining the 76ers. Next to a slew of underwhelming free agency options, he chose a team that will present him with young and old personalities to manage, on-court chemistry questions to answer, genuine Finals upside, as well as some wonderful downside potential in a city that’s internationally renowned for booing. We hit all of it on Greatest of All Talk: first with an emergency episode that we recorded two weeks ago (you can hear our disbelief an hour after the news broke), and then with a longer, 45-minute discussion this week . Two weeks later, I’m still shocked we’re here, and thrilled as a basketball podcaster. — AS Meta Earnings, Meta’s Timing Problems, The Financial Tail — Meta’s earnings were a bit disappointing; future promises about AI products were more disconcerting. Microsoft Earnings, Microsoft vs. Meta, The Efficiency Payoff — Microsoft’s earnings were compelling because they showed a clarity of strategy, lower costs, and a tangibility of application. The reason why is scarier. Google Earnings, The Frontier Case, Amazon Earnings — Google’s earnings seemed to confirm the Anthropic hedge; it was Andy Jassy who explained why their — and Amazon’s — capex was justifiable. Vibe-Porting and Meta Enterprise OpenAI Responds ADSL Made the Modern Internet Possible A Memory-Maker Makes History; New Robot Rules; The Open Weights Debate Rages; End of July News and Notes Six More Questions on LeBron in Philly, Revisiting the 2016 Draft, The Top 5 Dinosaurs Microsoft’s Plan for Platform Survival, Meta and the Market’s Permission, A Lack of Situational Awareness

0 views
ava's blog 2 days ago

letting an LLM write an indie web parody post

There's certain topics and writing styles that do really well on the indie web; especially on Bearblog's Trending page. There are enough joking and seriously annoyed posts about it every now and then. We cannot stop talking about how tech is bad, analog media is good, blogging is good, and how good RSS is. I was wondering whether LLMs are trained on enough data from here to generate a post that would do really well, and is realistic, while also parodying us bloggers a little. So I prompted: Generate a short to medium sized stereotypical post that could be posted on bearblog.dev that could easily reach Trending, to parody how always the same topics get recycled in the indie web. Make it a little satirical/over the top, but still realistic, just enough to let the viewer know it is a joke, but still enough that it is something someone would really post. And it delivered. Behold. A few weeks ago I archived my notes, closed my 137 browser tabs, and deleted the app that was supposed to organize the apps that organized my life. Nothing happened. The sun still came up. My coffee still got cold because I was busy reading something interesting instead of optimizing my morning routine. I've started carrying a small notebook again. It's imperfect. The pages wrinkle. My handwriting is terrible. But somehow that feels... human? The modern web keeps asking us to produce, optimize, engage, and monetize. The old web just asked us to make a weird page about something we cared about. That's why I like small websites. No algorithm. No growth strategy. No personal brand. Just people quietly existing on the internet. Maybe the future isn't another platform. Maybe it's RSS. Maybe the best software is software that gets out of your way. Maybe the best note-taking app is the one you stop thinking about. Anyway, thanks for reading. If this resonated with you, I'd love to hear your thoughts on your own little corner of the web. A little sidenote at the end roasted us: The parody works because nearly every paragraph contains a familiar indie-web trope: deleting productivity tools, rediscovering notebooks, criticizing algorithms, romanticizing the old web, ending with RSS, making broad philosophical claims from a minor lifestyle change, and closing with a gentle invitation for conversation. making broad philosophical claims from a minor lifestyle change Well, that answered my question. And this marks my first and only (partially) AI-generated post on my blog. Don't worry, I think that will be my last. If you want to check out my stance on AI, here is my tag for it . Published 07 Aug, 2026

0 views

The Beginning Place

Hugh, suddenly overwhelmed by the dreary, lifeless routine of work-television-work, runs from the house he shares with his mother. He does not know where he is going. He runs down streets and past suburban homes and finds himself at a strange and beautiful creek. The water is cool and revives him, brings him back to himself. He soon visits this place as often as he can, finding in it something sacred and unspeakable. He thinks he is the only one who comes there, until one day he sees a young woman, who likewise thinks this is her place. In fact, she has known of this place for much longer, and knows also of the strange people who live a ways up the trail. She knows, too, of a great fear that is terrorizing them, and that there is something—perhaps someone—they are waiting for. There are several allegories at work here, not the least of which seems to me to have some of Beowulf to it. But I think the book may also be seen as a kind of coming of age story, if coming of age was the work not of one but two, not of solitude or isolation but of connection, conjoining, coupling. View this post on the web , reply via email , or become a supporter .

0 views
Unsung 2 days ago

Seeing like a state

The post about the dark mode toggle reminded me of two similar things rattling in my brain. On the positive side, here’s a delightful interaction from macOS. I can easily maximize the window to take up half the screen, but the moment I start dragging it, it recalls and nicely restores itself to its original size: macOS designers correctly figured out that the window being maximized or half-maximized is a state – but it has to be a state dressed up as a size. The button entry point is the “state” version. But on the way in, there is also a more natural “size” version: you can have the window snap and maximize to half screen when you drag it to the right edge. And on the way out? You just saw it. You don’t have to switch the state to “non maximized” first, and you don’t have to restore to the original size by hand. Here’s a bad example – one of the macOS’s horrible settings pages: = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/seeing-like-a-state/2.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/seeing-like-a-state/2.1600w.avif" type="image/avif"> So far, it seems good. Some of the toggles are on, some off. You not only see a position of the switch change, but also the track under the switch is a different color to help you disambiguate. Nice. But now look what happens when I toggle off the second option, which the third and fourth option rely on: = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/seeing-like-a-state/3.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/seeing-like-a-state/3.1600w.avif" type="image/avif"> Processing this dialog visually, does it look like “on, off, disabled off, disabled on,” or does it look like “four toggles, each one inexplicably with a different shade of gray”? There are many solutions here: some visual, some IA, some systemic. Also, I use the graphite accent color, which somewhat exacerbates the issue, although it’s there with any accent color. But I wonder if one of the challenges here is that someone thought it’s important to show the state of the toggle even if it’s disabled, and everything else followed from that. This feels similar to the dark mode essay in that there will always be someone making that argument, and that argument will always feel stronger, because it will feel like it’s backed by logic. The system will make sense as a diagram. Each of its parts will come from a logical conclusion. So did the tri-state dark mode toggle . Or the Power/​Sleep/Wake keyboard buttons. Or Abort, Retry, Fail in DOS. Arguments for systemic completeness are always going to be easier to make than arguments for thoughtful simplicity. I sketched two possible solutions. They’re not the best ones, and you might recoil at them, since either one is a compromise. But that’s the point. = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/seeing-like-a-state/4.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/seeing-like-a-state/4.1600w.avif" type="image/avif"> = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/seeing-like-a-state/5.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/seeing-like-a-state/5.1600w.avif" type="image/avif"> #complexity #interface design #system design

0 views
David Bushell 2 days ago

They don’t make ’em like Sublime Text anymore

After a multi-year detour wading through the quagmire of bloatware that is modern software I’m back to coding in Sublime Text . I am finally at peace! Do you know how relaxing it is without the daily “update available” and “view release notes” ritual? There is nothing I dread more than release notes. Between VS Code and Zed, I must have glanced at a million lines of sloppy commit summaries. Can I remember seeing anything of value? Nothing comes to mind… Modern software sucks. It’s built by self-proclaimed “engineers” who make their chat box addiction your problem and still use Twitter. When I mentioned Sublime Text on the socials people went all nostalgic on me. Sublime Text is still around. The last stable build was released May 2025. There are dev builds as recent as July this year. I don’t know who develops it and I don’t care. All I know is that I can place my text cursor and read the surrounding code without a bombardment of popovers, pop-unders, pop-left-and-rights, pop-inlines, pop-in-and-out-too-fast-to-sees. It’s been a few weeks since I returned to Sublime Text. It does everything I need. There is not a single “feature” I’ve missed from VS Code or Zed. Many I’m glad to see the back of. The perfect dev stack is a collection of software that each does one job and doesn’t suffer main character syndrome. I want to code, I’m not looking to make lifestyle choices. I don’t want a bloated everything app. Don’t get me started on the “unified toolchain” plague! Show me the latest VC-backed build tool and I’ll show you ten lines of PHP that does a better job. If you’re fed up of the absolute state of things, Sublime Text still works. Thanks for reading! Follow me on Mastodon and Bluesky . Subscribe to my Blog and Notes or Combined feeds.

0 views
Unsung 2 days ago

“Solving a largely imaginary user goal”

On her blog, Lea Verou makes a case that each user-facing website dark-mode toggle should only ever show two options , but in a smart way. The challenge is that any dark mode toggle needs to actually accommodate three options: dark, light, and the default “whatever the system says” (which can be always dark, always light, or change with the time of day ). Many toggles simply pass that complexity onto the user: I want to get something out of the way: I don’t think Verou’s article as an article is fully successful. I feel like it spends a great amount of words to explain something not entirely as complex, and even the interactive playgrounds felt slightly too rigid and altogether confusing. If you care about (interactive) explainers, it might be an interesting case study in and of itself. But I am very much much onboard with the proposal and the line of thinking it represents. Verou suggests a “smart” dual state toggle, which still allows the website to follow the system, but shoves the complexity of the “whatever the system says” branch into the crevices between visible UI. Here’s how I understand it: This toggle will feel compromised, and you might immediately find some rare use case it doesn’t fully support – maybe attached to an imaginary user, or even an internal user. But Verou is absolutely correct in her insistence to fight through that: Tri-state toggles are implementation-driven UI. One of the most common UX mistakes is designing UI around the underlying data model instead of user goals. Good interfaces abstract away the underlying model and expose a model that aligns with user goals (unless of course these happen to coincide, which is rare). Now, it’s just a dark mode toggle. It might not seem like a difference between a smart dual state toggle and an explicit tri-state toggle is that much. But: A similar example might be that of PC keyboards in the late 1990s, which also exposed system complexity and pestered people with Power/​Sleep/Wake keys: = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/solving-a-largely-imaginary-user-goal/3.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/solving-a-largely-imaginary-user-goal/3.1600w.avif" type="image/avif"> Computers do not do that anymore, simply having a smarter singular power button, piped to a more sophisticated logic underneath. #complexity #dark mode #keyboard #web The smart toggle only has two options: light and dark. Mechanically, clicking or tapping the toggle brings you to the opposite option. Simple. If your new option is the opposite of system (e.g. you switch the page to dark mode if your system is in light mode), it will stay in that theme forever, no matter what the system does in the future. If your new option is one that currently matches the system, it will then continue following the system in perpetuity (e.g. it’s back to the default behaviour). “Whatever the system says” is not just one extra option. It’s also one extra weird option. It doesn’t feel like the other two. It’s seemingly repetitive. It’s often unclear what it does before clicking. It’s not obvious where to put it in order. Verou doesn’t mention this in her post, but even just seeing the word System next to Light and Dark feels complicated. (Auto is slightly better.) The cognitive load here might be larger than it seems. What is an interface if not a collection of a million challenges, each one seemingly insignificant on its own? Trivial things add up. One compromise here and one cheap decision there, and soon you’re talking real money. Thinking deeply about something like this gives one practice of dealing with complexity elsewhere, and facing even more difficult challenges where the stakes are higher and the compromises larger.

0 views
Jeff Geerling 2 days ago

I'm excited for Intel after testing the XPS 13

Shortly after Apple launched the budget MacBook Neo , Dell announced their response, a new low-end XPS 13 . Matching the Neo's current pricing, it starts at $699, or $599 with an educational discount. That discount is currently set to expire on November 2, and with the current component pricing insanity, I'd be surprised if we don't see a price increase on both laptops by next year. I ran the XPS 13 through my gauntlet of benchmarks , and published a review on my YouTube channel:

0 views
Kev Quirk 2 days ago

📝 2026-08-07 09:22: Follow up from this post - I just signed into my Vinted account (to delete...

Follow up from this post - I just signed into my Vinted account (to delete it) and it didn't send me an MFA SMS so how the fuck is me giving them my number securing my account in any way? The internet is fucked. Thanks for reading this post via RSS. RSS is ace, and so are you. ❤️ You can reply to this post by email , or leave a comment .

0 views
Kev Quirk 2 days ago

📝 2026-08-07 09:10: Brilliant. I'm such a "valued customer" that they couldn't even be arsed to put my...

Brilliant. I'm such a "valued customer" that they couldn't even be arsed to put my name in the email. Why is it even necessary to have a "business intelligence database" with a 3rd party? I ordered a fucking laptop. Can a tech company not create their own customer database? Full email here - https://cdn.kevquirk.com/framework-breach-email.pdf Thanks for reading this post via RSS. RSS is ace, and so are you. ❤️ You can reply to this post by email , or leave a comment .

0 views