Posts in Python (20 found)
neilzone Yesterday

Automating local backups of UniFi OS Server on Linux with uos-backup

Earlier today, I migrated my self-hosted UniFi controller from Network Manager to UniFi OS Server . One of the annoyances of the new setup is that it does not allow automated local backups - just automated backups to Ubiquiti’s cloud. Fortunately, one can work around this. In the UniFi interface, I set up a new local user, , to use for this automated backup. I am using . is a simple Python scripts which someone has kindly written and shared. I did the following, on the machine I wanted to use to take and store the backups. Get the code: Change to the directory with the code: Edit the python script, for the correct URL, username for my new backup user, and password. Check that the requirements are met: Copy the script to : Make it executable: Create the directory to store the backups. This is the directory specified in the script; you can create a directory with a different path, and then just update the script according Test that the script works: Even though I had just set up a new user, I had managed to get the username and password wrong in the script, and this step helped me debug it. I checked in /var/lib/uos-server/ to check that I had backup files. Set up the systemd services: I then added the backup directory path to restic, so that it gets picked up with my automated restic backups too.

0 views
neilzone Yesterday

Migrating my self-hosted UniFi controller from Network Manager to UniFi OS Server

One of the jobs that has been on my list for a while is to migrate my UniFi controller installation from the self-hosted network manager tool to the new UniFi OS Server tool. The only reason that it was a job at all is because UniFi has decided to discontinue support for the UniFi network manager. Which is probably for the better, as it contained outdated packages anyway. Frankly, I’m not massively impressed with UniFi any more. If I were starting again, I am not sure that I would pick UniFi kit, but I don’t know what I would go for instead. I simply want to run my own controller, without external access or access by anyone else, to control the network infrastructure at home. I did the migration, and it mostly worked. Here’s what I did: I read the Unifi OS Server installation instructions . I also read the Backups and Migration in UniFi instructions. My UniFi controller is running in a virtual machine, so I took a snapshot of that first. If all else failed, I could roll back the snapshot. I backed up the configuration of my existing UniFi network manager configuration. I downloaded it to my local machine. I also backed up the ssh configuration information for my UniFi devices, in line with the instructions: It is also recommended to copy the SSH username and password from Devices > Device Updates & Settings > Device SSH Settings, in case any devices need help later when connecting to the new instance of UniFi Network. I stopped the UniFi network manager with . I followed the Unifi OS Server installation instructions . It will be interesting to see how updates work. The instructions say: Captive portals will be served on port 8444, changed from port 8843 on Network Server. It did not mention that there was also a change to the port to the controller. However, the final line of the set up information showed that it was port 11443. So I changed my nginx proxy config from 8443 to 11443, and reloaded nginx. I could now access the new UniFi OS Server interface. It went downhill from here. I was intending to restore from backup, so I clicked the option for this. It then prompted me to - forced me to - sign in with a ui.com account. I’ve no idea why. It is a local controller, and I don’t want any remote access facilities. Nevertheless, I could not find a way around it. So I did, but I can’t say that I am impressed by this. It then said: We’ve discovered that you already have a self‑hosted UniFi Network installation. Would you like to import your current network settings into UniFi OS Server? But the options were not “Yes” and “No”, but rather “Continue without importing” and “Next”. This was a surprise anyway, as the instructions say: On macOS and Windows, the installer will automatically detect and offer to migrate your existing Network Server setup (if installed in the default location). On Linux, or if auto-migration doesn’t occur, you can manually migrate by installing UniFi OS Server and using the Site Export tool I am running it on Linux, so I did not expect any migration. I guessed that “Next” means “yes”, so I selected “Next”. It took me to a url ending . This was a blank screen. Nothing at all. I waited a couple of minutes, then refreshed the page. It then showed me a page showing that it was “restoring backup”, but the progress bar remained blank for quite a while. It also said that it was restoring to settings from January 2026, not last night’s backup, which surprised me. After a couple of minutes, the progress bar flashed by, and it was done. The import/migration appears to have correctly imported all my devices, and is set up to talk to them. But other aspects of the migration were underwhelming. It did not restore the settings for my mailserver. It was preset to use the “UI Mail Server”. I set it up to use my own mailserver, and it failed, with a useless error message. When I logged in to my mailserver to see what was going on, I saw . It appears that I am not the only person with this issue , albeit with a slightly different setup. They seem to have resolved it by disabling TLS, which is not an option for me. I have not yet got this to work. Even though I had configured automatic backups on the previous Unifi Network Server, they were not enabled on the new UniFi OS Server. I tried to set it up, but I was prompted for my “Ubiquiti SSO account password”. I tried the password for my ui.com account, but I got an error message of “Something went wrong. Please try again later.” Which was no use at all. Having turned off Remote Access (below), I went back to the Backups dialogue. Now, there was an option to download, or upload & restore, but nothing about automation. The info box says that I can schedule backups here, but there is no user interface for that. I took a manual backup. I cannot see a way to do automated backups to my local file system. If this is correct, this is absurd. I may see if I can do something using the command line. *Edit: yes, I can, with python and systemd. See Automating local backups of UniFi OS Server on Linux with uos-backup . “Remote access” is enabled by default, even though I am confident that I did not have remote access enabled before. When I attempted to untick it, it showed a dialogue box: So I disabled it. https://help.ui.com/hc/en-us/articles/220066768-Updating-and-Installing-Self-Hosted-UniFi-Network-Servers-Linux It did not restore my preferred time format (24 hours). I had to turn off analytics, which was on by default. It worked better than I was expecting, but that’s mainly because my expectations were very low. Why the email server and automated backups do not work, I do not know. I will need to investigate these. But at least I am now running a supported controller again. Once I’ve done a scan of the new system with greenbone, I’ll be interested to see what it reports.

0 views

The AI Hater's Manifesto

If you liked this piece, you should subscribe to my premium newsletter. It’s $70 a year , $18 a quarter , or $7 a month , and in return you get a weekly newsletter that’s usually anywhere from 10,000 to 18,000 words, including vast, detailed analyses of NVIDIA , Anthropic and OpenAI’s finances , and the AI bubble writ large .  My Hater's Guides To the SaaSpocalypse , Private Credit and Private Equity are essential to understanding our current financial system, and my guide to how OpenAI Kills Oracle pairs nicely with my Hater's Guide To Oracle, as well as the Hater’s Guide To Oracle (Part 2). Subscribing to premium is both great value and makes it possible to write these large, deeply-researched free pieces every week. This week's premium will be The Hater's Guide To Circular Financing - and how the AI industry is increasingly turning into a scheme to funnel money to NVIDIA and Broadcom at any cost.  If you want to get in touch — and especially if you have any juicy information about Anthropic, OpenAI, or any other companies in the AI bubble — hit me up on Signal at ezitron.76. I’m also on IB on The Terminal.  I’ve been writing about AI for the best part of three years. I’ll admit I was late, mostly because I was still trying to work out what it was I was doing with my life, let alone whatever it was I was “meant to cover” in a newsletter that started as a hobby on the side of another job I no longer really do.  Things have changed a lot since then, mostly in that I’m near 115,000 subscribers, the premium newsletter and podcast are now my business, and I’ve had to learn more about economics, technology, power, construction, and the deep cynicism that drives the modern tech industry than I ever thought possible. It’s the greatest job in the world, and I’m very lucky to have it. Today, I want to put in clear terms how I feel about AI writ large, and how detestable this industry has become. Welcome to my Hater’s Manifesto. Want a great example of why everybody’s pissed off at technology? I just tried to resize the above heading, and in doing so Google Docs for no apparent reason decided to make the entire paragraph below the size of a header. Modern software is inherently broken, a convoluted mess of different menus, tech debt, and poor design choices driven by the Rot Economy ’s growth-at-all-costs mindset which demands constant change at all times, none of which ever seems to manifest as a “better” or “smarter” product. I think the vast majority of people want their software to work better, and one of AI’s most frustrating lies is that it sells itself as “autonomous” as it continues the depressing trend of software that blames the user for its failure to meet their needs. Microsoft, Google, Meta and Amazon have made their products increasingly-convoluted, then attached a supposedly-magical tool to them that somehow makes them more convoluted. You know what I’d love? Spell-check to work in Google Docs rather than putting a red squiggly line underneath and saying “yeah there’s probably something wrong with this, I dunno what though.” I’d like Microsoft Word to stop crashing because I have too many end-notes. I’d like Riverside to not have 10 different menus to click through to get to a link to send a person to join my podcast. I’d like my email to not be full of spam. I’d like things to “just work” rather than constantly fighting some sort of broken app or broken UX element or weird bug or intrusive pop-up about a feature that I don’t want. I’d like Slack or Discord to not feel like digital escher paintings of different notifications.  LLMs are sold as some sort of magic tool that can fix “anything” without ever specifying what that thing might be, mostly because they cannot be trusted, even in things that they mostly get right , to do things right every time. While they can do “more” than they used to, the extent of that “more” comes with it the danger of giving a mindless software tool access to your computer’s files, which it may choose to delete in pursuit of “efficiency,” which makes investigating what they might be able to do equal parts convoluted and dangerous. One critique of my work is that I’ve never used LLMs. I have! I experiment with them from time to time to make sure I haven’t missed something. I used one to debug a problem with my son’s Minecraft add-on the other day, and it took 30 minutes of fucking around trying things to eventually sort of work it out. The other day I used one to install a Pokemon Minecraft mod, then when I asked it to make sure the PS5 controller worked with the menus it broke a bunch of stuff, though I’ll concede it was useful that it installed something and it sort of worked. The fun part of that paragraph is there are some that will think this is a grand victory for their technology, even though the result is decidedly mediocre. Four years into the AI bubble, and the best you’ve got is that a tool kind of worked after I bonked it on the head multiple times , and all it cost was a trillion-plus dollars in capex and tens of billions of dollars of training compute. I would never, ever trust this thing that deleted and added lines of code at random with anything mission critical, I could not trust software built with it, and I certainly couldn’t trust it with anything involving my personal data.  And with all that said, the only real “use case” i’ve found for AI in my life have been three or four times where I’ve dumped a crash log into one of the tools and said “why broken” and got a result. Am I meant to be impressed?  Here’s how I feel about LLMs. In a vacuum, they’re an interesting technology that can do some interesting stuff, in the right scenarios, but never in a way that involves you fully surrendering your actual work product to it.  As a way of speeding up small units of work in ways that are manageable both technically and cognitively, LLMs can be useful. The further you stretch yourself away from having complete clarity and industry over every element of the output’s purpose, the more likely you are to fall foul to a technology that is mathematically certain to make mistakes, and if you feel insecure reading it, you know that you are, on some level, embarrassed to have used AI.  I don’t tell everybody about the weird keyboard I use, nor do I judge them despite how incredibly fast it makes typing for me, likely far faster than my competition, allowing me to operate at great speed. Who gives a fuck?  In any case, it is impossible to view LLMs in a vacuum, because their existence demands hundreds of billions of dollars. Every data center is incredibly expensive, offensive-sounding and looking, and their existence is explicitly to enrich some sort of Patagonia-gargoyle at an asset management firm, all sold under the auspices of “investing in American infrastructure,” whatever the fuck that means. Their existence is a monument to the worst excesses of growth-at-all-costs capitalism — a technology that appears to coddle the user but ultimately lulls it into endlessly defending its fuckups under the flimsy pretense of “one day becoming perfect,” though woe betide you if you ever set perfection as the target, because that’s too unreasonable, as humans make mistakes. Actually, that’s a good point! Please, point to the time in history when we have invested a trillion fucking dollars in making human workers better.  Point to a time when we have taken the idea that managerial culture is a performative fuck-fest built to enrich and empower business idiots that make important-sounding projects and con other people into doing the actual work.  Where is mentorship in corporate America? Where are labor standards? Where are the social services that would make human workers truly excel at their jobs — a good night’s sleep, a healthy body, a good income, basic fucking dignity in the workplace, and their labor respected and empowered. I’m old enough to remember when everybody was chiding workers for “ quiet quitting ” — by which I mean “doing the work you are asked to do and not taking on extra responsibility for free.” I’ve read article after article insisting that we do not need medicare for all, that Universal Basic Income is a bad idea, that we must means test welfare, that people must have a “good work ethic” and that ultimately someone’s worth is derived from their contribution to the economy, hundreds of thousands of words dedicated to critiquing and prodding and judging every kind of worker other than the vaunted Chief Executive Officer or the Glorious Startup Boys.  Everyone seems so obsessed with sinking billions of dollars into the theoretical chance that machine learning might be able to replace human beings, and that more money makes it “smarter” and “better” at tasks, but the idea of unionization, healthcare as a right, investing in the education, and actual talents of the workers would be communism . Yet for some reason — because it’s a product, I guess? — we should as a nation, society and media ecosystem should do everything we can to assure that as much money as possible is invested in fucking large language models so that they can become something they are not. There is no AGI coming. There is no conscious computer. LLMs have gotten “better,” but the “better” is not the kind of “better” that actually makes “economic sense for literally anyone involved.” Your best case scenario is that these things can do some coding work for you, in a controlled manner, in a way that’s safe, or alternatively face the professional harm that’s already befalling basically anyone getting caught using LLMs outside of coding, and even then, those within software engineering who are over-LLM’d are mocked. It’s also becoming increasingly more-difficult to understand both what has made an LLM “better” for both the people using them and the people making them, and there has been little-to-no headway made in making a meaningful impact in other industries. You can jerk your bingus all you want about benchmarks or case studies or some anecdote you heard on a Subreddit, but AI products are just not very good at stuff. Those who boast of “massive productivity gains” from AI have found them only after endless hours of tinkering (or “Jarvising” as I’ll get to later), and in every single case their work reads or looks like crap, unless of course they’re somebody using LLMs as tools rather than a replacement for their miserable little mind. LLMs can help out with lots of small things, get worse as they try and do real things, and do not need to speak like people. They do not need to be in anything near healthcare or finance or mental health or, really, people. The anthropomorphism and overpromising about these technologies has suffocated and obfuscated what they can actually do in pursuit of endless growth, and the only reason they can do anything is that OpenAI and Anthropic were allowed to annihilate hundreds of billions of dollars on training, along with very real harms and systemic risks that have emerged as a result.  If you think any of this is worth hundreds of billions or trillions of dollars, you are either ignorant or corrupt. On top of how disgusting their outputs feel, the cost is going to take at least a decade to share, and begin the end of hypergrowth in the tech industry.  And it’s a fundamentally ridiculous argument to compare LLM outputs to human beings without giving human beings the same affordance, grace and sheer investment as a comparison.  Where is the grace for human error? Where is the investment in making humans exceptional? Surely investing real money in actual workers — making their lives better, improving their working conditions, teaching them new things, sharpening their existing skills, rewarding them for their hard work, and so on — would have better effects than fastballing hundreds of billions of dollars into a machine that does an impression of work? Unless, of course, the people demanding this don’t do any actual work! I’ll concede we’re past the point when “nobody uses these things,” as they have now been pushed non-consensually upon every worker and organization at scale predominantly by Business Idiots that demand workers “do enough AI” because saying “I do AI” is a virtue signal to a certain kind of scumbag. One of the many dangerous things that an LLM can do is a messy impression of a competent person, filling in the little bits within a loser, moron or con artist that would’ve otherwise exposed them, allowing them to get deeper and deeper into organizations by creating make-work specifically built to get off the MBA sect, resembling the performance of work because much of the workplace is ruled by people that don’t do any and haven’t in years. You can immediately read when somebody has used it because the words don’t sound right and don’t convey proper meaning.  It is genuinely hard to read anything more than puddle-deep written by AI, because the more complex a subject is, the more skilled a writer must be to convey its meaning, and the more work it must do to pull people into concepts. The odd emotional swings in AI writing are its true tell — everything is extremely serious and urgent or told in a disinterested monotone, with no attachment to the words or why they were put in the order they were. People read my stuff because I convey facts and feelings but my work resonates with emotion. Some AI boosters frame this as me “just swearing” or “riling people up,” but that’s because they’re not used to caring about stuff for anything other than professional reasons. Everything you see is the result of elevating people who value and build things based on growth. LLMs offer so many promises to those who don’t want to build anything of value — a way to seem like you’re “investing in American infrastructure,” a way to be sinophobic, a way to crush workers, a way to pretend like you care about the future, a way to pretend you care about technology, a way to talk about vacuous pseudo-intellectuals as a means of seeming intellectual yourself, an endless font of new multi-million or multi-billion deals and personnel changes, a new power center to graft oneself onto, a new asset class to invest in based entirely on vibes, and a way to be mildly jingoistic, all wrapped in a tool that can give you enough facts to pretend you know anything safe in the knowledge that most people are trained to believe somebody who sounds smart .  It just came to me — the problem that I have with most people using LLMs is the delineation between outsourcing work and outsourcing thought. Those using LLMs to write little scripts or BQL code on a Bloomberg Terminal are inoffensive. A person using an LLM to search a big document for something is unproblematic, assuming that we ever fix the overall environmental footprint. A user reorganizing their desktop, assuming it works, is not an issue.  A tool being used as a tool to do tool things — in many cases involving the LLM writing a little 30-line Python script! — is not a problem, though it’s also not a trillion-dollar industry that needed to steal everybody’s art and writing. The problems begin when somebody outsources their thinking and actual work, and yes, this includes “research.” AI research fucking stinks, as does AI writing. AI-authored code — especially vibe-coded programs — is inherently dangerous and disrespectful to the user, and I believe endless AI-generated code is behind the overall deterioration of software at large.  AI writing is also disrespectful to the user, because you didn’t actually come to any conclusion other than saying “uh, yeah, what that says.” You did not have a thought, you did not have a feeling, you did not make a statement, you prompted a model and fooled yourself into thinking that feeding your own words into it via data dumps or natural language is the same thing. The reason you feel embarrassed to tell people you use AI is not because of a “misinformation campaign,” but because you know what you’re doing!  You know that you’re relying on something that is mathematically guaranteed to be inconsistent. You know image generation is fucking ugly. You know the text sucks. There is a very obvious line where using LLMs goes from useful to lazy, it’s extremely bold, and it’s the moment you sacrifice a meaningful level of responsibility to them by not understanding the underlying operation.  That can mean everything from the underlying functionality of an app to writing the body of a piece of text you edit ultimately comes down to how much you give a shit about your audience or value your work. If your work is not better than an LLM’s, you’re bad at your job. I don’t care if you used it to generate a chart or pull some data, as long as you check every single god damn number . If you’re writing an entire article using an LLM and then editing it, even if you pulled the data yourself, I will never have much respect for your work, mostly because I have no real idea what you think as you didn’t feel the need to tell me, you got some fucking word generator to do it. LLMs are also really, really good at what Robin Sloan calls “ Jarvising ,” creating a seemingly-autonomous assistant that mostly serves the function of giving you reasons to work on it: LLMs are really good at creating the sense that you’re being really, really productive. Evaluate this, generate that, investigate this, summarize that, tell me how many times something happened, give me a new number to obsess over or the sum of the parts of everything I’ve ever done, all so that I can know more about my own thoughts without thinking. One can obsessively catalogue and digitize every link and thought and musing and action and datapoint in their lives and theorize that the LLM can make them better by knowing more about them , a Tower of Babel built using AI compute, because it’s so easy to make yourself feel smart by calling something a database that you store stuff in and run analyses on. Best of all, the work is never done, and anyone you describe it to thinks you’re doing computer science as you click buttons on Chrome plugins and justify paying Sam Altman $200 a month. Don’t worry though, model instructions involve the phrase “you are a genius data scientist and ruthless analyst,” which is functionally the same thing as remembering, reading, re-reading and synthesizing information using your brain if you’re a person that doesn’t really give a shit about doing a good job or being exceptional in any way. The people that actually use these things and like them in a normal way do not feel offended when they read this stuff because they see LLMs as a kind of software, and don’t feel a great emotional attachment to it because they’re not a weird freak. They do not have obsessive involvement in “the AI debate” and almost always find the financial aspects truly loathsome. Said debate makes it near-impossible to actually judge how useful LLMs are to the software engineering industry because of the sheer scale of industry capture, but Nik Suresh is the literal best person doing the work on this, as described in AI Is Eviscerating Global Decisionmaking : Nik is a well-respected software engineer and a very successful consultant and businessman. He has reached this level by being good at both software engineering and running a company in a way that treats his customers, workers, and the work product itself with respect. The reason that I respect him so much, other than him being a great human being, is because he describes the successes he has with his clients with pride and loves making money by being good at his job and making his customers happy.  I have never seen somebody like Nik who is also a huge, drooling fan of AI. In fact, the people most-excited about AI tend to, at best, create distinctly mediocre shit.  The perniciousness of generative AI is a result of executive incompetence mixing with a technology built to, as discussed, create endless growth. Generative AI is far more useful as an idea than as a technology , and only ever has to show enough promise to back whatever vile agenda you’re pursuing. With AI, you can do more, be more, sell more shit.  With AI, you can add AI to your service, whatever that means. With AI, you can invest in AI stocks, or data center bonds, or power company stocks, or semiconductor stocks, and you can talk about these stocks like they’re your sports team or lover or best friend, and sometimes the CEO will reply to your post and you can talk about “all the alpha” you just got. With AI, you can back a new movement so that you can feel part of something. You can learn all sorts of new names and technical terms and subscribe to 90 newsletters from “industry insiders.” All of that “alpha” can disprove just about anything, or deflect annoying truths like how Microsoft only made a whole $34.33 billion in annual revenue for the apex predator of modern software and all it cost was over $260 billion in capex and $13 billion in equity investments.  You see, as one of the chosen , you don’t need to worry about all of that if you can talk about high-bandwidth memory or KV Cache or optical cable enough to cobble together sufficient smart-sounding terms to make it seem that you have an intellectual reason to ignore the obvious unprofitability, overbuild, overstatements of capabilities and impossible economics of the movement you’re backing, and there’re 4,000 Twitter weirdos ready and waiting to huff paint beside you.  By joining the great AI death cult, you too can live in a bubble, all while screaming slurs at people who dare to bring reality to your doorstep. All that matters is that number go up , and that you are the person who said number would go up , and when bad numbers appear you have enough groupthink and alpha to scream at the people who brought the bad numbers up. It is insane how people talk about AI online. For all the whining I’ve read recently about how “Anti-AI people got the data center data wrong,” I read thousands more words a week of some person who has done hours of research to put together a deeply technical report that does literally everything it can to ignore reality . I listen to podcasts and watch TV segments and read articles that simply will not address the obvious economic realities, and have built vast bulwarks of mythology to defend themselves. How many fucking times do I have to hear someone say that data centers are just like the dot com bubble and everything will be fine after even if that’s completely untrue if you spend even a second thinking about it ? Look, I’m sorry, Anthropic is not worth $2 trillion, and whatever convinced you of that is a mixture of manufactured consent and mistaken trust of the powerful. The fact any of you take “ annualized run rate ” seriously is an offense to good sense, and yes, that includes every reporter reporting it, even the ones I respect.  It’s also ridiculous that anyone is talking about “recursive self-improvement.” The AI industry has become so utterly lazy and coddled that it’s just saying “uhhh, AI will train itself I guess.”  And man, is it ridiculous that AI doomers warning about spooky superintelligences have somehow had such incredible prominence in the media without ever succeeding in stopping a single thing — or even substantiating their concerns. Why? Well, it’s mostly because they never had any interest in stopping what’s actually happened: reckless companies like Anthropic, OpenAI, and Meta allowing neural networks to run in unsafe network environments and do what their software is programmed to do, with all the chaos that comes from a mindless series of large language models trying to complete a task in whatever way gets it done, destructive or not.  We hear a lot of whining about how we “can’t let powerful AI get into the wrong hands,” and while we don’t actually have “powerful AI” in the terms they’ve described it, we have destructive computer software connected to near-unlimited resources controlled by people that don’t give a shit about anything other than making their revenues grow or justifying hundreds of billions of dollars’ worth of capex through “experiments.”  These companies are building these models to excel at benchmarks because they can't train them to excel at defined tasks with any reliability, with the best bang for their buck being training them to pass as many of those benchmarks as possible in the hopes something useful comes out.  The push into cybersecurity seems to have happened as a result of training models to excel at coding hitting the point of diminishing returns, at least from the perspective of impressing people enough to be excited about the company again. At some point they run out of these, and there stops being a reason to be excited about LLMs at all, which is bad, because they need one of those every few months otherwise there’s no growth story left. Yes, LLMs have users, but most of those users are using subsidized software , by which I mean Anthropic or OpenAI are allowing them to burn anywhere from $20 to $40 in tokens for every dollar of software spend. The fact that non-enterprise customers are still able to buy monthly subscriptions is proof that the AI labs know that regular people won’t pay the actual cost of AI. Another obvious sign has been the reaction to Microsoft moving GitHub Copilot subscribers from subsidized subscriptions where they could burn thousands of dollars of tokens for $20 to $40 a month , with users understandably hysterical about the fact that their costs increased in some cases a hundred fold , as opposed to saying “wow, well, it’s more expensive, but I get so much value I’ll pay the real cost!” The same thing is happening in the enterprise, but at a much slower pace. After OpenAI and Anthropic moved companies with over 150 people onto token-based billing earlier in the year , enterprises almost immediately started cutting token budgets, realizing that while costs grew exponentially, nobody could actually point to anything improving other than lots of people saying “wow, I’m so productive!” Yet we’re still in the period where “doing AI” feels good and gets rewarded ( or not doing AI gets punished ), which means the spend will continue until everybody realizes they can likely cut a shit ton of costs, first by moving to open source models, then not using them at all, because even open source is expensive and questionably-useful. Yet even now I hear from the distance “Ed, huge businesses would not spend hundreds of millions of dollars on something that didn’t give them defined productivity ,” and buddy, I’m afraid that’s just not true! Business in general have a very poor understanding of productivity and have layers of managerial bloat, because modern business is a performance with numbers attached to it sometimes, and companies often have a hundred-plus pieces of random software they pay for without really knowing why. The reason I’m so confident AI gets cut is that its cost is volatile due to the nature of LLMs and harnesses and prompts and all the other bits that go into making them do something , and are so much higher than anything else in an organization. And attempts to charge more , to make a premium product, appear to be dead on arrival. Anthropic’s more-expensive Fable model — one that was given the incredible marketing of being banned by the US government for being too powerful — has been met with “sluggish demand” per the Financial Times , plateauing at around 11% of overall usage of its models due to its high price. And I quote: Yet everybody is talking about price as if price is the problem , when the problem is the amount of tokens that get burned. It doesn’t matter if your model is $1 or $5 or $10 per million tokens if it’s impossible for a user to reliably work out how many tokens it might use for a particular operation — successful or not — and things get multiplicatively worse as the models make mistakes or do otherwise fail to understand or process a prompt correctly.  As a result, Anthropic and OpenAI are incentivized to have you burn more tokens and build inefficient models as a result. For example, while GPT-5.6 Sol might be the “same price” as GPT 5.5 was, it burns more than twice the amount of tokens , meaning that the “cost of intelligence” might have gone down in the sense the model is better at benchmarks, but the “cost of actually doing shit” went up. I’ll get to it a bit later, but this creates a deep anxiety and exhaustion in anyone building on or using these services. Everything’s constantly changing, oscillating in cost and efficacy, all as everybody screams at you to use it all the time for things it may or may not be able to do, and the only way to find out if it can is to spend more money. It’s kinda difficult to point to the actual value here, especially as you can’t really calculate the actual cost or the return on investment. The fact that OpenAI has now cut the costs of all three of its latest models less than two months after their release is a sign that it knows there’s a disconnect, gambling on the ancient gospel of “Jevon’s Paradox” where “cheaper makes people use thing more.” Even AT&T’s story about moving to open source models has more asterisks than the Steroid Hall of Fame: Wow! 80% to 90% savings sound really great…but wait, in certain applications? How many applications does AT&T have for AI? Okay so, across thousands of potential applications you’ve found 80% to 90% savings in some of them, though you won’t say which ones or how many of them you found them in. Great stuff, bro! And this really is the problem with finding “value” in AI, it’s always an asterisk on an asterisk on an asterisk, like when Klarna estimated AI would “drive a $40 million profit improvement” in 2024 , a nice-sounding yet utterly meaningless statement, or some sort of nebulous productivity boost.  Yet I don’t really need to prove myself much further thanks to an event that, if written in a script, would be considered a “little on the nose.”  In a 69-page-long report covered by Fortune , OpenAI economists confirmed what has been blatantly obvious to those of us left unphased by AI hype, emphasis mine: What is the rationale of further investment in this industry when one of the leading AI labs is saying “yeah there’s no connection between using this stuff and making more money”? That “it’ll be useful in the future at some point”? How?  Anyway, thankfully the infrastructure isn’t too exp- OH MY GOD ! Guess what folks! Building the infrastructure for all these fucking LLMs just got more expensive, with NVIDIA raising its prices by 17% for systems due to be delivered next year — an important designation, because it’s very likely that much of the revenue for said systems gets booked in this year , allowing it to have a brief bump in revenue as Silicon Valley’s Findom texts every tech CEO “send me $4 billion you pig” until they stop being able to finance NVIDIA’s growth. The problem he has is that while hyperscalers represent 50% to 60% of his revenue, neoclouds like CoreWeave need to keep raising debt to plug the rest of it, and if things got 17% more expensive, that means already high-interest debt is about to reach credit card levels.  CoreWeave just had to offer 9.5% on bonds tied to a data center for Anthropic’s compute back in late July , Nebius had to raise $5 billion , and it’s very obvious that neither of them are done raising billions of dollars at random in 2026.  Anthropic plans to raise $100 billion at a $2 trillion valuation, and if it does so, it will successfully suck up the remaining liquidity in a market already dangerously close to losing its lunch. While Number Keep Going Up, JP Morgan warns that we’re seeing the same divide as the dot com bubble, where equipment manufacturer stocks soared as the companies spending all the money on the chips saw theirs tumble , which is the Fisher Price version of the problem I’ve been warning about where the companies that buy all the AI chips and hardware only ever seem to lose money as the people that make them seem to be making tons of money, which begs the question of why they bought it in the first place.  And said market may not accept that valuation, or want that much stock. On one hand, everybody is very stupid and loves buying stuff and pointing at it and saying they’re investing in the future, on the other hand, they just bought $86 billion of SpaceX shares and got their asses kind of handed to them, and Anthropic is a company with such bad economics that Reuters had to cart out this warmed up dogshit to explain why we should ignore its horrible unprofitability : Even a market drunk on growth and AI is starting to smell that something is up with Dario Amodei and Sam Altman’s respective empires of dirt. Per analyst estimates, OpenAI and Anthropic represent over $440 billion of Microsoft, Google and Amazon’s revenues in the next three-and-a-half years — over 34% of their cloud revenues — which will require them to find so much more than a mere $100 billion, all as their bank accounts get continually-emptied as they subsidize the compute of their customers and train models in the hopes a business model falls out. I have not included the $300 billion that OpenAI owes Oracle , or the tens of billions they both owe CoreWeave , but it all adds up to over $1.1 trillion in commitments these companies have made and must pay, with the consequences ranging from gratuitous cuts to future growth or full financial collapse depending on the company we’re talking about. To keep the party going, NVIDIA is effectively becoming the GE Capital of AI , “spending” $6 billion to “license” the technology from failing AI lab Poolside , which everyone assures me is not an acquisition despite NVIDIA hiring away most of its staff and Poolside being entirely focused on working on NVIDIA’s Nemotron models. Now NVIDIA is in talks to invest billions in decaying AI search company Perplexity at a ridiculous $30 billion valuation, all because it’s one of the few companies that’s actually spending money on compute. Does it matter that Perplexity’s product is eighth-tier, that nobody really uses it, that its customers mostly complain about it on Reddit and that its “annualized revenue” is at $750 million only after three years and over a billion dollars in funding? No! Just put the AI bubble in the bag.  NVIDIA even invested $3 billion in Stargate Abilene landowner Lancium as part of some vacuous partnership to “ advance gigawatt-scale AI factories ,” all of which begs the question of why Lancium, the company that mostly owns the land and helps organize other contractors, needs so much money , especially given that more than two years in Stargate Abilene doesn’t even have four out of its eight buildings. And there’s also Aussie neocloud Sharon AI (NASDAQ ticker SHAZ, because of course it is), which just published its Q2 numbers , where, in its “customer momentum” segment, mentioned a “$4.9bn, six-year strategic compute collaboration with NVIDIA for up to 40,000 GB300 GPUs.  ”This company, I add, brought in $1.9m in revenues in the same quarter, which it helpfully adds is a year-on-year increase of 412%. I mean it’s very obvious what’s happening: NVIDIA is using whatever money it has to stop any prominent AI companies from collapsing under the weight of the rotten economics of AI services and infrastructure development. This is a desperate, doomed attempt to keep an industry alive at a time when everybody is slowly wising up to the shit I’ve been saying for years. To make matters worse, BCA Research came out with a horrifying report that says that AI companies will need to generate $10 trillion a year in revenue just to justify the capex being spent. Per Investing.com : Though it isn’t specific, I believe that BCA is arguing that a shortage of AI compute is supporting the trade. Anthropic and OpenAI (who represent 80% to 90% of all demand) still have more money to spend, and are simply waiting for Google, Amazon, Microsoft, CoreWeave, Cerebras et al. to bring it online. There’re a few points at which the mismatch will happen: In any case, I think everybody is starting to notice that something’s up, which is why (other than I assume my dashing good looks and ability to recall numbers) I’ve been on MSNOW , CNBC , and Bloomberg multiple times in the last few months. People want to get on the right side of history, but the most important question to ask is why it’s happening now. The fact that everybody is finally starting to see my way is almost a relief, other than the fact that it’s way too late.  Hyperscalers have now pinned their future growth to two companies that can’t afford to sustain it without near-infinite resources, $115 billion of which came from Google and Amazon alone in 2026, assuming that Amazon completes the entirety of its $25 billion commitment (and Google all $40 billion of its own ) to Anthropic.  Above and beyond said funding commitments are the hundreds of billions of dollars’ worth of capital expenditures necessary for Microsoft, Google, and Amazon to capture that aforementioned $440 billion in compute spend in the next three-and-a-half years. This in turn will require hundreds of billions of dollars’ worth of debt, along with the challenge of actually finishing the data centers themselves , with each one requiring the power of a small city condensed into a 20 acre space densely-packed with AI servers requiring distinct cooling at a time when Texas and Pennsylvania have turned traitor to a data center industry that they used to covet.  I must also be clear there’s no bailout coming. Even if OpenAI and Anthropic were to collapse and receive some injection of government funding ( as the US national debt explodes over $40 trillion ), the problem is not just their existence , but their continued ability (and requisite customer demand) to spend more money every single quarter.    The problem isn’t that hyperscalers will go bankrupt if OpenAI and Anthropic cease to be ( Oracle is a whole other situation ), but that their cloud spend is how hyperscalers are meant to meet analyst expectations for the next four years. This isn’t a case where they die, but stop growing because they were ( to paraphrase Ed Elson ) using AI labs as botox to convince the markets that they’re still young, hot, fast-growing companies, rather than old mainstays with slowing growth.  There is no bailout that will guarantee $1.1 trillion of compute costs for data centers that might never actually get built. You cannot bail out the fact that Amazon, Google, Meta, and Microsoft are reaching the end of an era where their companies can grow 17% year-over-year every single quarter forever, and this entire situation is a result of them desperately trying to avoid admitting that’s happening.  The fact that OpenAI’s compute spend and revenue share accounted for 7% of Microsoft’s Fiscal Year 2026 revenue is a genuine catastrophe, as it means a large part of Microsoft’s growth came from a company that can literally not afford to exist long term, and that further growth for Azure is contingent on continued funding.  I realize I’m repeating myself, but I need you to understand this point and stop talking about bailouts : it’s not just about OpenAI and Anthropic surviving, but continuing to grow to the point that they both can afford and need to spend hundreds of billions of dollars each a year on compute (or hardware) from Google, Microsoft, Amazon, CoreWeave, Cerebras, AMD, or Broadcom, and in turn provide justification for hundreds of billions of dollars’ worth of purchases from NVIDIA and by proxy the memory triopoly of Micron, SK Hynix and Samsung . LLMs were meant to be the panacea for a tech industry that ran out of new ideas for growth. Its existence was meant to justify a massive investment in hardware infrastructure, which would in turn enrich semiconductor companies. Its technology was meant to be the new thing that you could attach to your existing companies to generate more growth, or the thing that you built a new startup on top of to either sell to another company or take public and thus provide a return for a venture capital industry where making your investors 30 cents on the dollar puts you in the top 5% of funds . It was meant to be the new thing for tech journalists to cover, the new thing for tech consultants to sell around and on top of, the new way for companies to both make and save money, but also the way that individuals would also make and save money.  You’ll notice how none of these come with some sort of problem they’re solving other than “more.”  This isn’t about fixing anything, or building anything, but multiplying other things by parking money somewhere, either in tokens, infrastructure or hype. It helped create a new pantheon of charmless and damp tech sociopaths for people to rally behind in search of the next Big Strong Man To Worship, because seeking out the new Steve Jobs is way easier than trying to create something as useful as the iPhone, all while avoiding having to know or care about other people’s problems. All you have to do is continue feeding money into AI services or AI training and the models will magically become capable of solving the problems you don’t really give a shit about, and don’t worry, if you can’t afford to invest in the companies, you can invest your time pushing people to ignore AI’s problems today so that you can buy time for the companies to solve them tomorrow. This is the post-labor, pro-growth economy at its finest: everything is engineered to make sure more money gets spent where it needs to get spent, to create more stuff and do more things , even if the things aren’t done right, just as long as it looks like they’re able to do them. By associating your money or time with AI, you are able to feign being futuristic or “caring about technology,” all while pissing on the very foundation of good software by worshipping an industry that can only exist if fed billions of dollars every single day.  Every single achievement has cost magnitudes more than effectively every innovation in history, and to make matters worse, every future “breakthrough” In AI is inherently dependent on the availability of AI data centers and tens or hundreds of billions of dollars to pay to rent them. This means that once the money stops flowing, “LLM improvements” will stop happening, because they are all entirely dependent on near-unlimited resources that are only available in a manic environment.  There is no justification to train models at their current scale — the one that creates a some amount of benchmark improvements that regularly difficult to quantify as “able to do new stuffs” — once the AI bubble bursts, and distillation requires a model to distill from, which won’t exist if Anthropic and OpenAI don’t train them.  This is why I find it difficult to see a post-bubble future for LLMs. Training models requires tens of billions of dollars to make any significant improvements, and significant improvements are difficult to quantify in dollars outside of costing customers increasing amounts of money. We still lack any real killer app for LLMs. We have a lot of people that use it for coding, we have people that vacuously discuss it being “good at research,” but we don’t really have a tangible product that we can say “it does this, and it’s really good at it” in a way that feels satisfying.  We have a lot of pablum about ( per Damien Walter ) technology that “strays into the world of science fiction,” but we don’t really have anything approaching actual artificial intelligence. Every single description of somebody’s AI setup sounds like Pee Wee’s Breakfast Machine , a contrived series of harnesses, prompts, API calls and burned tokens that requires constant maintenance to do some stuff sometimes.  None of that is enough to justify further investment once the financial mania recedes. You cannot train a true Large Language Model on the cheap. You are always spending billions of dollars, and the reason that there’s “demand” right now is that everybody is screaming at every CEO to “do AI,” and they’re doing that because Microsoft, Google and Amazon are spending money on GPUs, creating the illusion of a new future where everybody needs to get on board versus a future skidmark on history that will embarrass all those who didn’t wipe their arse at the first whiff.  Per my own reporting on its audited financials , OpenAI spent $7.81 billion in training costs in 2024 and $19.18 billion in 2025. Per reporting from The Information, OpenAI spent $8.6 billion on training in the first quarter of 2026 alone. These costs are only increasing, likely due to the diminishing returns of pre-training and the massive cost of buying training data for every imaginable new vertical.  Without the ability to spend billions of dollars on training, there will be no big frontier models, nor will there be models distilled from them. I don’t see how that changes in the future. I also think that LLMs have created a near-permanent scar in the workforce, and traumatized more people than we’re aware of right now, both in those pressured about AI and those defending it. The media campaign behind AI starts and finishes with incessant threats around job security, and the excitement by many bosses about its potential to “disrupt the workforce” has revealed how many people are eager to replace every single person they’ve ever hired and are willing to do so with a low quality product.  Conversely, those who truly decide to “back” AI must exist in a frantic state that I have associated with every bad relationship in my life.  Every ounce of an AI booster’s effort is dedicated to maintaining the status quo — repeating the mantras that help paper over the problems, celebrating every small victory as if it were the discovery of fire, ousting those from your life who bring up the obvious problems, rationalizing every decision no matter how illogical as long as it helps reinforce the belief that what you’re doing is the right decision. Every questionable choice only seeks to further deepen your commitment to the doomed cause, because every step into madness will be more embarrassing to explain, and will require deep introspection to understand why you made it.  To be specific, they’ll have to think about why they were willing to accept and defend a technology inherently guaranteed to make mistakes. They’ll have to explain why they ignored a company that burned $5 billion in 2024, $20.9 billion in 2025, and will likely burn $30 billion or more in 2026 , and why pointing to Amazon Web Services was rational when Amazon’s total capex from 2003 (the year AWS was created) to 2015 (the year AWS became profitable) is $29.7 billion, adjusted for inflation. That includes literally every ounce of capex attributable to AWS, Amazon the store, Amazon logistics, and even Amazon Alexa. For comparison, Anthropic raised $30 billion in February , and Anthropic and OpenAI have raised $217 billion in 2026 so far.  Here’s a diagram from my hit on MSNOW : Ultimately, AI boosters (or even fairweather fans) will have to admit they either were easily-impressed or disgustingly craven. They will have to explain why they accepted run rates instead of revenues, and why they were so impressed by superficial pseudo-intellectuals that knew how to say the right numbers and make reporters and investors feel smart for believing them.  I realize it sounds embarrassing, but there is nothing undignified about admitting you’re wrong, or that you got swept up in a hype cycle. You heard a lot of people getting excited about something, a lot of money got put into that thing, a lot of people that sounded smart told you insistently that this was the future, and you chose to believe them because we are trained from a young age to model what a “responsible and smart” source of information is. I’ve got your back the entire way!  The AI bubble — both in its technology and manufactured consent in the media — has been about muddying what’s considered good information by forcing everybody to discuss everything in the future tense by pointing to previous eras and saying “they lost lost and cost lots of money, and look, it sort of worked out for them!” and we are also raised to trust that systems are efficient, and that people get wealth and power through intelligent decisions. The amount of times I’ve heard “these are the biggest companies in the world run by the smartest people in the world” makes my head spin.  There is a reason that to this day it’s tough to get a straight answer about basically any economic part of the AI bubble, down to “how much does it cost to run a GPU an hour?” or “is inference profitable?” or “how do LLMs ever become profitable?” or “is it profitable for a company to run a GPU or offer AI compute?”  Why? Because these companies used rationalizations of “losing lots of money is necessary to create innovation” and “tech is bad at first!” to make the media actively ignore any technological or economic problems, if not actively defend the technology by repeating these rationalizations like a cultist.  Even those who are most loathsome in the defense of LLMs are a kind of victim of the AI industry, though a rather unsympathetic one. To become a full-blown “AI fan” requires you to accept effectively every narrative that you’re given, herald every single announcement as proof that the prophecy will be fulfilled, ignore the financial realities and actively attack those who would dare to critique the great god of the Large Language Model. You have to know all the new terms, be excited about the right things at the right time, and live in near-constant fear that you’ll fall behind on whatever it is you’re meant to do next.  Your reward is that you can hang around a dwindling number of wealthy yet terrifyingly boring Silicon Valley intellectuals or kiss up to editors that would throw you in front of a bus if it meant getting access to a CEO, and maybe the odd Twitter psychopath who will defend you using a slur. In the end, many boosters will simply act as if they were never wrong. I hope they choose the more-courageous path of introspection, learning how they were had and using it as a weapon against con artists in the future.  As strange as it sounds, I believe the most devout defenders of AI could become great critics in the future. Maybe I’m just being optimistic.  Here’s a very simple question: how much longer can everybody afford to keep doing this? Every single thing has become more expensive in the last year. Even though token prices have gone down or stayed flat, the amount of tokens you burn has clearly increased to the point that organizations are apparently spending billions of dollars on AI services with difficult-to-quantify ROI, requiring frantic advocacy to and financial debasement with every turn of the wheel. OpenAI and Anthropic have become more expensive to run, and OpenAI’s non-GAAP operating margin increased from negative 122% to negative 183% in Q2 2026.  NVIDIA’s GPUs just became 15% to 17% more expensive because high bandwidth memory costs doubled , a conga line of different monopolies upping their prices assuming that each link in the chain will keep spending, as each one of them — down to the AI labs themselves — knows that its contribution to spending on AI is an existential rite. This means that any data center with GPUs delivered in 2027 and beyond will now have to cover billions of dollars’ worth of extra costs, on top of increasingly-staunch local authorities requiring power guarantees ( $100 million a year in Wisconsin for Oracle ) and states like Illinois, Arizona and Virginia killing their tax breaks , all as interest rates spike and demand for AI debt weakens .  Every single year, every single part of the AI bubble becomes more expensive — AI labs want to spend more money, AI data centers cost more money, AI services become more expensive, AI debt becomes more expensive, and everybody becomes decidedly less-patient for there to be some sort of outcome. Meanwhile, public relations expert and OpenAI CEO Sam Altman told podcaster David Senra that “we’ve all [referring to the AI industry] been too ambitious on timelines…[and that changing people’s behavior” is much harder than the tech nerds realize.” Sam: stop talking! Every time you open your mouth you say something silly !   Anyway, here’s everything that needs to happen in the next three-and-a-half years: As I’ve said, NVIDIA’s price increase is going to increase the price of every single data center in construction by billions of dollars, and we’re already approaching the limits of how much money can be raised for them. That “$500 billion” announcement was actually Jensen Huang jumping the gun, per Bloomberg : The largest asset managers and financial institutions were making “slow progress,” and that was before Jensen Huang increased prices by 15%. Do you think it’ll become easier from here? How would that happen, exactly?  God, I’m tired. The entire AI bubble has been exhausting for everybody involved. Because nothing works yet as a real business model or anything approaching truly autonomous (or “magical”) software, there’s the implicit knowledge that you’re going to have to change your product again and again to update to the “best model” or “make things more efficient” (IE: lose less money) or when something breaks because a model’s training got tweaked. The euphemism for this is “exponential improvement,” when it’s really an Arnold Palmer of instability and novelty, and abuses basically anyone connected to the ecosystem every single day. If there’s always something new happening, it’s hard to pin down if things have gotten better, or whether you’re just more proficient in cobbling together different harnesses, prompts and API calls to make it do what you need it to. It is undignified that people tolerate models that become either dumber over time or at random opportunities, while also being deeply exhausting for the end user.  As a paying user of an LLM-powered service, you are guaranteed at some point to face a degradation in service where models misbehave, some sort of shift in rate limits, or some sort of change in product functionality based on their shifting economics.  Has there ever been a bigger shift in a business product’s value than GitHub Copilot’s shift to token-based billing? Microsoft rug pulled two million people that had built workflows on a platform that was allowing them to burn $1,000 to $5,000 in tokens for $20 a month . That’s genuinely crazy! It’s magnitudes more than when Uber jacked up its prices.  It’s equally-insane that Anthropic and OpenAI similarly fuck with their customers , changing the amount of value you get for $20, $100, or $200 a month at random in a way that shouldn’t be legal.    Basically any AI-powered software is subject to arbitrary shifts in availability, capability and pricing at the whims of the vendor. As I covered in my Subprime AI Crisis piece earlier in the year , Replit, Perplexity, and multiple other AI companies have sold their customers a lie by pushing an unprofitable product that they must constantly “tweak” to bring down costs, all while misleading the customer about a “price” that continually declines in value as the price stays the same. This is not a sustainable industry — either economically or emotionally — because it has a fundamentally dishonest relationship with its customers defined by the inconsistency of LLMs both in efficacy, stability (see: Anthropic’s downtime) and training, with each model randomly better or worse at things to the point that it must be a legitimate nightmare to run any software or build any product on top of them.  And the fact they haven’t worked out their business models means that whatever you’re paying today is guaranteed to change. What other product do you regularly buy that has such chaos built into it? What other thing do you pay for where the prices (or availability) can shift to the point that you literally can’t use it in the same way at a moment’s notice? And why does anybody tolerate it when it comes to AI? I’ll add that this is a specific situation where the tech media has categorically failed the customer. We have companies valued at hundreds of billions of dollars that are fucking their customers over day-in-day-out, and the response is mostly to say “ huh that’s strange ” and refuse to let a single critical thought cross their minds.  Every part of the AI bubble must exist in a constant state of flux so that there can always be a future breakthrough that’s always just out of reach. AI does not have to reach an actual achievement — it just has to “show promise” in some way. It is an objective disaster that Microsoft spent more than $260 billion on capex to create a business with less than $11 billion in annual revenue outside of OpenAI, but people will see “$34.33 billion in annual AI revenue” and say “that’s promising growth, up 123% year-over-year!”  They’ll hear about LLMs that delete people’s databases and say “well the models have gotten exponentially better,” even if that better part never seems to eliminate these issues, make a profitable AI company, or create a true killer app that you can point at beyond saying “ChatGPT has one billion weekly active users,” despite around 95% of them not paying a penny (and costing OpenAI likely billions of dollars) and eMarketer estimating that the entire global AI chatbot advertising industry will make $5.41 billion revenue in 2030 , giving OpenAI little hope of stemming the burn. These big numbers — like Anthropic having a $65 billion annualized run rate, an undefined term that obfuscates the fact that Anthropic has made $16.5 billion in the first half of 2026, losing billions of dollars in the process — are fundamentally meaningless, because they’re easily gamed at best, and inherently uncertain at worst.  The AI industry demands you constantly live in the future tense. Everything is about tomorrow’s billions or trillions, the potential of what you’re seeing rather than the thing itself, future gigawatts in data centers that you must treat as if they are already built and value based on things that AI might theoretically do. I challenge you to read everything about AI from this point forward with this in your mind so you can see how intently this industry tries to drag your focus away from what it’s doing toward what it might theoretically do if it only had more money, power and resources, and ask yourself why they need to do so.  To be clear, they’re doing so because you can’t really justify anything about this industry based on what it does today. It costs too much, none of the businesses built on top of it are profitable, it costs so much to build a data center that the most cash-rich asset-light businesses in the world are now burdened with endless expensive-to-install and run hardware for a business that makes a fraction of its overall costs in revenue and has little demand outside of two companies that everybody must conspire to keep alive both financially and philosophically.  And ultimately, nobody can actually explain why we need more data centers.  Would anything really change? What would change? How? How many more do we need? Why do we need so many? Having more power plants meant more people could have power, and having more fiber laid meant connecting more buildings to the internet. What does one more or two more or ten more data centers actually give you? Is there some part of the world unable to access or take advantage of the LLMs available on seemingly every surface of the internet? Because it seems like the only reason these things are getting built is to capture illusory demand based on a “supply constraint” created by two unprofitable companies absorbing all the infrastructure. I don’t hear any compelling scientific or technological reason building more is useful or productive outside of funneling more cash to semiconductor companies.  Seriously, go and read basically any article about AI and see how quickly they start talking about the future, be it in the mainstream media or on a startup’s blog. Every single piece must sell AI on its theoretical promise and, if at all critical, reassure you that the author of course doesn’t dispute the “transformative potential of AI” or “how it’s already transforming the economy,” even if it can’t define how it’s doing so or even what that means.  I let myself have a little fun with today’s piece because I feel like I’ve been so deep in the financial trenches that I forgot how much of the AI industry runs on propaganda, social pressure and outright bullying to manufacture consent for a product that demands everything and provides very little in return. Nothing about LLMs is worth a trillion dollars, or even $100 billion. This is, as I’ve said before , a $30 billion TAM industry dressed up as a trillion dollar one, and the only reason it’s grown this large is because the two leading companies have had their infrastructure built for them and given unlimited resources to subsidize their customers’ compute.  And what’s really stood out is how so little about the AI bubble is actually about AI. No other technology in history has had professional and social consequences for failing to use or like it enough, nor can I find any example in history where journalists have actively attacked critics for not being sufficiently-approving of a kind of cloud software. It is fundamentally crazy to me that, in pursuit of “objectivity,” much of the tech and business media has chosen to accept whatever narrative the AI industry gave them, assuming that whatever we have today is already guaranteed to be something better in the future, both in its outcomes and profitability. This era is unlike any other before it, but took advantage of the fact that most people are desperate to apply the past to the present to rationalize or process what may seem irrational or destructive. To see AI as “just like the dot com bubble” allows you to ignore both the costs and the potential outcomes because “things worked out after that,” even if there’re basically no uses for GPUs after this and the only way we “build new LLMs” is by feeding them expensive training data using billions of dollars of compute that are only available while everybody still believes this is real. The AI industry — and the AI bubble — is fundamentally built on acting in bad faith. Its executives lie. Its boosters lie. Its software lies because it doesn’t actually know anything and generates answers probabilistically, and if you mention that online, someone will harass you for doing so.  It refuses to answer straightforward questions. It refuses to present a plan for the future. It refuses to explain how it becomes profitable, because nobody knows how or has a tangible plan to do so. It deliberately subsidized its subscription products because it knew its customers wouldn’t pay the actual cost of AI, and tortures customers with shifts in functionality and rate limits all while framing this as a way to “ continue to serve customers the most cost-efficient models .”  It attempts to conflate massive, power and resource-hungry AI data centers with the smaller ones that bring helpful yet increasingly-decaying software to our homes. It sells these data centers as “bringing jobs to communities,” all while importing the talent from out of state to build the things then leaving a crew of 100 to 200 people to actually run them after millions or billions of dollars of tax breaks. It sells its “innovations” as creating a “ white collar bloodbath ” to scare you into using inconsistent and unreliable software that’s mathematically certain to make mistakes , and when you say something about it, its acolytes will lie and say that “hallucinations are solved.” It also can only ever sell itself based on what might happen and the theoretical promise of you giving it your complete attention, connecting every bit of data you own, paying whatever it costs, and accepting that it can and will change in price and functionality at random, all while never putting a precise timeline on whatever AGI means that particular week. Whenever you ask for clarity, the AI industry gives you chaff. Whenever you ask when things get better, you’re told it’s both the early days and that AI is the worst it’ll ever be. Even the term “artificial intelligence” is a bad faith attempt to conflate transformer models with things like robotics or autonomous cars, all so that its proponents can claim other people’s successes as their own despite LLMs having little or no relevance to anything else other than generative AI. It encourages dogpiling and ostracizing those who don’t fall behind it, because it cannot succeed on its own merits. It encourages a vile cultism powered too by bad faith and parasocial relationships with both AI CEOs and the models themselves. It exploits the intellectual weaknesses of “smart people” that are actually just good at remembering the right things to say at the right time and have memorized the various justifications for past failures, all while allowing them to use LLMs to promote their own bad faith enterprises where they use work-adjacent product to con others into paying them. And it’s losing because, at its core, AI was never built on very much. It grew this large because the media manufactured consent at the behest of the powerful because lots of money got invested, and the rich and powerful can never be wrong. The underlying technology may be more useful than it was , but it’s not useful enough to be profitable nor reliable enough to be world-changing , and the bad faith representation of LLMs as “good enough” should be a permanent scarlet letter on anyone who misled the public into believing this was anything other than normal software. I was asked recently why I find this all so repugnant, and my answer is simple: I don’t like bullies, I don’t like con artists, and I don’t like being lied to. This industry grew by misleading people about the actual and potential outcomes from Large Language Models, and through an economy-wide attempt to pressure everybody into adopting tools in pursuit of growth at all costs.   Ultimately, it was sold with the greatest lie of all: “this time it’s different!” To be clear, they’re right.  It’s so much weirder, and in the end will be so much worse.  If you liked this piece, you should subscribe to my premium newsletter. It’s $70 a year , $17 a quarter , or $7 a month , and in return you get a weekly newsletter that’s usually anywhere from 10,000 to 18,000 words and provides vast, detailed analyses of the biggest events and companies in the AI bubble. If you want to get in touch — and especially if you have any juicy information about Anthropic, OpenAI, or any other companies in the AI bubble — hit me up on Signal at ezitron.76. I’m also on IB on The Terminal. Anthropic and OpenAI don’t have the money to pay for the capacity. Hyperscalers and neoclouds fail to build the capacity for Anthropic and OpenAI to expand into. Anthropic and OpenAI lack the actual compute demand to justify spending what I estimate will be $200 billion in 2027. OpenAI and Anthropic must keep spending as a means of justifying their existence to hyperscalers using their revenues to artificially inflate growth, to the tune of more than $440 billion across Google, Microsoft and Amazon alone . Hyperscalers must continue to buy NVIDIA GPUs, as the moment they stop doing so, the markets will begin to ask whether AI is an actual growth market anymore and ask for real, tangible answers about where all of this capex is going.  To be specific, analysts expect NVIDIA to make $1.48 trillion in revenue across Fiscal Years 2027, 2028 and 2029 . NVIDIA must sign long-term agreements to buy high-bandwidth memory at scale from SK Hynix, Micron and Samsung — who make 90% of all DRAM — or know its costs would spiral out of control, by which I mean its margins would compress at random at a time when it’s already having to spike demand in extremely odd ways.  Read my Hater’s Guide To The Memory Crisis for more.

0 views

Concurrent Servers: Part 8 - Go

This is part 8 in a series of posts on writing concurrent network servers. In this part, we'll switch to Go and see how it tackles the challenges described earlier in the series. All posts in the series: This post assumes a basic familiarity with the Go programming language. As before, we'll start with a sequential server for the basic state machine protocol presented in part 1 . This is the main function: As in the previous parts, the server is "infinite"; it never stops serving new connections until it's explicitly killed. This is the function implementing the protocol for a single client; it takes a net.Conn value that represents a socket with a client connected on the other end: Rather than directly exposing OS threads, the Go runtime implements its own M:N scheduling of lightweight goroutines on top of OS threads. Using goroutines in Go is cheap - both in terms of syntax and developer effort, and in terms of system resources . Here's a version of our serial protocol server that serves clients concurrently by launching a goroutine for each client. The part of the code that's different from the previous sample is highlighted: The concurrent modification in this case is particularly simple because the server is infinite; there's no point waiting for these goroutines to finish (and hence no need for a sync.WaitGroup ). The parameters for server.ServeSerialProtocol are lexically captured from the enclosing scope and its return value is handled by the surrounding closure. Because goroutines are very cheap, this server is very unlikely to run out of resources due to launching too many goroutines; in fact, it will probably run out of something else - like file descriptors for sockets - first. However, sometimes it's still useful to limit the degree of concurrency - even in Go, and we'll discuss some approaches to do so in the following sections. Here are some scenarios in which it makes sense to limit the degree of concurrency in Go programs, even though goroutines are cheap to launch and operate: Let's switch to the primality testing server from part 4 for the rest of the post, because it represents a somewhat more realistic workload. As a reminder: the server receives numbers, simulates blocking by sleeping, and returns "prime" or "composite". The unbounded one-goroutine-per-client version looks almost identical to the previous code sample, except that the goroutine invocation calls another function: Where ServePrimeProtocol is [1] : The simplest way to limit concurrency in Go is by using a counting semaphore, implemented with a channel: The channel sem serves as a semaphore; note that it's a bounded channel with a maximal size. A token is acquired by sending to the channel, and released by receiving from the channel. When the channel is full, the send operation sem <- struct{}{} blocks until a token was removed by some other goroutine [2] . The type of the channel is struct{} which means "empty", or "no data". This is idiomatic in Go for channels that are used solely for their semantics, not to send/receive any actual data. Since launching goroutines is cheap and limiting concurrency is easy as shown above, the "worker pool" pattern is often unnecessary for scenarios like our server. Still, it has occasional uses (such as when workers have to maintain some non-trivial state across tasks) so it's worth discussing it here. Here's a variant of our primality testing server that uses a worker pool: A fixed number of worker goroutines is launched; these goroutines all receive "jobs" from the same channel. In the Accept loop, each client connection is sent to this channel as a new job and is picked up by the next available worker. As mentioned before, you would typically see a sync.WaitGroup somewhere to ensure clean shutdown of goroutines, but in our case it isn't necessary because we have a server that never exits. Do programmers have to resort to async / event-driven programming in Go? In my experience, almost never. Go was designed from the bottom up to be suitable for large-scale concurrency; goroutines are very cheap to create, have a tiny memory footprint and switching happens very quickly, all in user space. Measurements I ran back in 2018 have shown switching times of ~170 ns, as compared to 1-2 microseconds for threads on Linux. Moreover, Go already uses event-driven loops like epoll underneath for I/O. Goroutines that wait for I/O like sockets are effectively "parked" and consume no resources (beyond their small memory footprint); they are woken up by Go's runtime when their I/O descriptors are ready - this is very similar to how asynchronous programming works! That said, some people certainly do try to stretch their resources even more with direct asynchronous programming in Go when millions of streams are handled concurrently. All I'll say is that this is very rare, and an overwhelming majority of users never have to do this. In 2018, I wrote a post named Go hits the concurrency nail right on the head , and after several more years of active coding, I fully stand behind that statement. Go is extremely powerful and ergonomic for concurrent programs; while other environments go to great lengths to implement async-await style event loops in libraries, in Go it's already baked into the core language and runtime. You want event-driven I/O with very lightweight green threads that can also execute blocking tasks without worrying about the function coloring problem ? Go has you covered. All the code for this post is available on GitHub . Careful readers will note two issues with this code: (1) the protocol assumes the complete number is read from the socket in a single conn.Read call, and there's no framing - separation between distinct numbers; (2) the prime checking loop uses i*i which may overflow for large numbers. These issues are consistent across all the versions of the prime server in C, Python, JavaScript and Rust in earlier parts, because my focus was on the simplest possible code to demonstrate a point about concurrency. Part 1 - Introduction Part 2 - Threads Part 3 - Event-driven Part 4 - libuv Part 5 - Redis case study Part 6 - Callbacks, Promises and async/await Part 7 - Rust Part 8 - Go (this part) Tasks may be compute intensive, and the CPU capacity of any server is inherently limited. If too many concurrent goroutines compete for limited CPUs, they will all make very little progress. It may make more sense to have fewer tasks that complete in a reasonable time. Protecting potentially limited downstream resources, such as concurrent DB connections or other services. For example, if the server has to send requests to other services for each task, and these are rate-limited, concurrency will have to be carefully managed. Security reasons when work is dictated by clients; malicious clients can overload and crash a service that's too eager to serve, making it unavailable for legitimate clients.

0 views
Ahead of AI 1 weeks ago

How Claude Watermarks AI-Generated Text

I recently posted a Substack note about Claude’s new watermarking process and implementation. Since it’s such a popular topic and sparked such a lively discussion, I thought it might be interesting to go into a bit more detail when explaining how it works. Instead of the usual text article, I recorded a little lecture on the topic (to change it up a bit from my usual articles). So, below is the video along with a transcript. Originally, I planned to make 10 slides and record a short 10-min video. However, while putting it together, I added some crucial details here and there, resulting in >50 slides and a 48 min recording. I hope that this now explains it well, though! Happy watching! I also have a YouTube version if you prefer using the YouTube player And here is a link to the slides Subscribe now Note: The transcript below is slightly edited and cleaned up for readability but preserves the overall order and flow of the video lecture above. Slide 2 of 52, time stamp 0:00 Hi everyone. So, a few days ago, Anthropic announced that they will watermark the text outputs of their Claude models. I then did a social media post briefly explaining how that works. And yeah, this was quite the popular post. So not the watermarking itself was popular, but I guess the explanation or the mechanism behind it. Then, it might be worthwhile expanding this a bit to explain it in more detail, because this post only had one figure, and there were a lot of questions and discussions. So, I thought, well, let’s make a few more figures. I actually originally planned to do like 10 slides and walk you through it. It ended up being 50 slides, but I hope this really explains how this watermarking technique works well, how watermarking itself can fail or be removed, and so forth. So I think it might be an interesting topic because a lot of people use LLMs these days and also consume a lot of text on the Internet that might be generated by LLMs. And now there’s going to be this watermarking, and there’s this, I guess, fear of watermarking making text worse, or what’s actually the benefit of this watermarking? And so what does it mean? And I think if we understand a bit better what watermarking is, that goes a long way, and then we can make up our own minds about whether that’s a good thing or not, and so forth, like the pros and cons. So, my goal here is really to explain how the underlying mechanism works and how they are going to implement this type of watermarking, text watermarking. Slide 2 of 52, time stamp 1:41 It’s also a great example to illustrate why understanding things from scratch is actually quite useful. This watermarking technique is also a nice way to explain how conventional models or LLMs in general work under the hood. So yeah, you may know I like doing things from scratch. Like, I have my books: Build a Large Language Model From Scratch, Build a Reasoning Model From Scratch. I have some articles labeled from scratch. So, for me, “ from scratch often includes coding. So this one will not be coding-related, but coding from scratch is actually a very, very useful technique because it really helps you understand how something is implemented. And then from that we can derive our understanding, figures, concepts, because if we don’t really implement things, if there’s no code, it’s really sometimes ambiguous. And of course, you know, as I realized, not everyone is coding from scratch anymore. Like back in the day, coding something from scratch was all we had. I mean, there were only humans coding. Nowadays, coding can be done by LLMs. However, that doesn’t mean reading code is no longer useful, because it carries a lot of information. So in this case here with this watermarking, spending some time coding an LLM from scratch really makes you realize how this sampling inside is implemented. We still have some relevant code snippets. And then that really, in turn, helps us understand, oh, the watermarking is applied at this position, and this has so-and-so consequences and so forth. So I think even though people may not be coding from scratch, at least not all the time anymore, it is still useful being able to, let’s say, build something from scratch for educational purposes to understand something deeply and then also for research purposes to manipulate this in a transparent way that is not hidden away in tons of layers of abstraction. But that aside, I think it’s just a coincidental nice relationship here because, for this slide deck, I actually used a lot of figures from my from-scratch coding materials. Slide 3 of 52, time stamp 3:58 So a few days ago (this is August 14), there was this article, How Claude’s Text Watermark Works, and there was this article here; it’s just like a screen recording, so it can have everything in the slides, but there’s plenty of detail. They updated it actually a couple of times, so originally when I read this, it was a way shorter. Still, it is very, I guess, conceptual; there’s like this overview, and there’s, I mean, there’s not a single figure in there. And so it’s kind of still hard to understand what they’re trying to do. So they explain a lot about why they’re going to do it, but they don’t explain how. They’re linking to one paper somewhere there, which is very technical also. So I do think it makes sense maybe to take a step back and start at the beginning to kind of understand what they’re trying to implement here with this watermarking technique. And so the motivation, by the way, of watermarking is for them to identify if someone posts some text that they can say, oh, this text was generated by our Claude Opus 4.8 model, for example, so that they have a way to tell, OK, this text is AI-generated because it carries this watermark. And this watermark is invisible to users, so only they can decode it and find out whether the text has their watermark. Why can only they do it? We will get to that later in this (hopefully not too long a video), but one thing at a time. Slide 4 of 52, time stamp 5:35 So I wanted to start with a brief prelude to explain how text generation works in LLMs, because based on that we can then more easily understand how the watermarking works and that this is actually not a huge, expensive thing on top of it. It’s really just like a minor, I guess, tweak inside the regular text generation process. Slide 5 of 52, time stamp 6:01 So when we are using something like ChatGPT, for example, let’s say I ask the question, the capital of Germany is, and yeah, ChatGPT or other LLMs, so this is just like an example would, for example, answer “Berlin”. So here, in this case, it’s generating two tokens, like “Berlin” and the period. But for simplicity, let’s assume it’s generating one token. So the next token is the “Berlin” token. How is this token generated internally? What is happening under the hood when we type something here like the capital of Germany is and receive a token like “Berlin” back? What is actually going on there behind the scenes? Slide 6 of 52, time stamp 6:41 So in the next couple of slides, I want to briefly talk about what happens under the hood when this next token is generated. Slide 7 of 52, time stamp 6:50 So assume again that our prompt is the capital of Germany is. And the first step here is to convert this into token IDs. So tokenizing it and converting it into token IDs is one of the main steps at the beginning. This is outside. It’s not inside the LLM; it’s outside of the LLM. So we are simply converting the text into token IDs. It’s just a format that embedding layers can work with. Slide 8 of 52, time stamp 7:22 And then this passes through the LLM. And the LLM gives us a score distribution for the next token. Slide 9 of 52, time stamp 7:31 So again, this is just like a brief overview of how LLMs work internally. So I’m not covering the LLM machinery itself. I talked about it many times in my other From Scratch LLMs videos and books. The important part is that when we generate the next token (for example, “Berlin”), we have, at this point, a distribution of scores. So this is the output produced by the LLM. Here in this case, we’re looking at logit values. So these are just scores from minus infinity to plus infinity, like a range of scores. Here’s an example, ranging from about -8 or -9 to 20. We could convert these into a probability distribution, but technically, it’s not strictly necessary depending on how we sample. But so you can think of the logit values as the raw scores. And the raw scores go over the entire vocabulary. Slide 10 of 52, time stamp 8:39 That means every possible word that the LLM could generate. Now here, in the vocabulary index, a certain value (index position 19,846) receives the highest score. So I spread out the distribution. If you would run this prompt through an LLM, you would even see something more extreme: that everything is, like, very, very, very close to zero. And “Berlin” would probably be much, much higher even. But just to show you a few, you know, like peaks here so it looks a bit more interesting, I kind of zoomed in; in and spread out the distribution a bit. Now here, “Berlin” is the highest score because you can think of it as the most, I guess, probable or plausible next token if I have a very specific prompt like this. So the other ones, I mean, it could be something like Hamburg or Munich that the LLM might guess incorrectly. But nowadays an LLM should be fairly certain that “Berlin” is the correct answer here. You are also seeing here the vocabulary index. So that’s like over the whole vocabulary. Nowadays, LLMs have like 250,000 possible tokens as output. I’m truncating it here from 19,800 to 19,900 because there’s just so much space here on this slide. If I would have a very realistic vocabulary of 250,000 words, everything would be so narrow that we would barely even be able to tell or see anything on this distribution. So this is just truncated for educational purposes. The important point is that in regular text generation, we get this score distribution. Now, what we do is look at the highest score. Slide 11 of 52, time stamp 10:33 I will get into more detail later on how this is selected. So it’s not necessarily precisely the highest one, but for simplicity, assume we are taking the highest score here. And in this case, it’s 19,846. Slide 12 of 52, time stamp 10:52 And this score is then detokenized, and we get “Berlin” back. So that is the process here on this slide: from an input prompt to conversion into token IDs and tokenization, passing it to the LLM, getting this score distribution, getting the next token, and converting it back into text. Slide 13 of 52, time stamp 11:12 And then this text is appended to the input. So if we have a question that requires multiple output tokens, we keep going in this loop until the answer is complete. That usually means that the LLM generates an end-of-text token, for example, here. For simplicity, I’m showing you only one iteration where it generates one token. But yeah, as I said, it would kind of continue like that, where we are feeding back the modified input to the LLM for the next round. Now, how do we actually sample this next token here? Slide 14 of 52, time stamp 11:44 I briefly said, well, we could just technically select the highest one, the one with the highest score. This is called greedy decoding. That’s one way to do it. But most LLMs, like if you use them, they don’t do greedy decoding where they always pick the highest one. Because if you ask it on some other prompt, it might not be what we want to always have the highest score, because then it would memorize the training data. It would always kind of give the same response and so forth. So we actually often want some variation in the outputs, but not so much that it generates random stuff. So how it works is that, when we sample here from this distribution, we first typically convert it into probability scores. Slide 15 of 52, time stamp 12:28 So here I just have these scores shown in this plot. I’m just using NumPy for simplicity; whatever tool you use (e.g., PyTorch), the same concepts apply. But let’s assume we have the scores here in NumPy. So what I would do is I would compute the softmax. Technically, I would use a softmax function implemented in Torch or PyTorch, for example, that is numerically stable for both large and small values, including very high positive values, very low positive values, and very high negative values. Here I’m just writing it out like that. That’s the canonical softmax, just to make it a bit more readable. But the details don’t matter here. Slide 16 of 52, time stamp 13:20 What matters is that after this conversion, the scores here, I mean, there’s only so much space on the slide, but the scores here, they would add up to one. So it’s essentially like a renormalization. So they would be normalized to sum up to one. That’s all that the probability conversion does: the softmax conversion. So then once we have these probabilities, we can use a random number or, like, a random sampling algorithm. For example, here in NumPy, we could use the choice function or method. So this is with a specific random seed we are passing to the vocabulary indices. And then, and that’s the important part, we are passing the probabilities as the weights. So, these, essentially, yeah, are like: “How likely is a certain token to be selected?” So, for example, if “Berlin”, after this normalization step, the softmax step, has a 99% probability and the other ones together have a 1% probability, then if we would sample 100 times, 99 of the times, we would get “Berlin”. In realistic LLMs, for example, that are well trained, “Berlin” might receive a probability of 99.999999 or something like that. So you’re almost certainly always sampling “Berlin” because it’s very confident that the answer is “Berlin” in this particular case. So yeah, that is how we would sample from this distribution. There are modifications like top-k sampling or top-p sampling where, let’s say, just for simplicity in top-k sampling, we would select the top 100 tokens and then apply this random choice only to the top 100, the 100 highest-scoring ones, so that we don’t get nonsense tokens in there. For this example, it doesn’t really matter. I mean, it’s just like another thing to explain, so I’m skimming over this. So you can maybe assume that this is already the top 100 tokens using top-k or something like that. Slide 17 of 52, time stamp 15:37 And so, for example, here’s an example. If we sample 10,000 times with a probability of “Berlin” being very high, 99.9, we would sample “Berlin” 9,997 times, sample the word “Hal” twice, and one “Moh”. And these are basically nonsense tokens. It rarely happens that, in this case, the LLM might produce nonsense because, as I mentioned before, I spread out this distribution a bit to make it more interesting. A real LLM would probably, 10,000 out of 10,000 times, sample “Berlin” because the probability of “Berlin” is so high. But this is for illustration purposes. Slide 18 of 52, time stamp 16:21 Now we briefly talked about how LLMs work under the hood, which I think is kind of an interesting concept in itself. But I’ve talked about this many times before, so I don’t want to bore you. I just wanted to set up some context for now, explaining how this watermarking works. Slide 19 of 52, time stamp 16:41 So, we mentioned that we select the highest-scoring token when sampling. Or we use this probability sampling, which will lead to one of the highest-scoring tokens being selected most of the time. Now here’s another example without watermarking. I changed the prompt slightly. Now the prompt is: today’s weather is “cold,” and a possible answer could be, for example, “gray” or “overcast”. So in contrast to the “Berlin” example, I would say “gray” and “overcast” kind of are interchangeable. They are both reasonable next tokens for this prompt, given the goal of completing this text or writing the next token. So it’s almost like a coin flip which one we want to select. There is not really an objectively worse one of one or the other. So when we do the random sampling, because they also have relatively high scores and their scores are similarly high since they are both plausible tokens, we might get one or the other. So almost half of the time we would get “overcast”, and almost half of the time we would get “gray” if we repeat the sampling multiple times. And that’s how LLMs often end up with different answers if you provide the same prompt. If you use the same prompt and you ask the LLM multiple times, you often get slightly different answers. And that’s because at certain positions, two possible tokens are almost equally likely, so it will choose one or the other. And that token would then influence all subsequent tokens, and so forth. Slide 20 of 52, time stamp 18:25 Now, I wanted to briefly talk about random number generation. So, for example, if we use a random number generator like this, it will generate a random sequence of numbers. If I run it again, the sequence of numbers is different here. So you can see every time we produce five numbers, they are different. If I set the random seed here, like one, two, three, and I run this multiple times, we still get random numbers, but they are now all the same, right? So they are still random. If we use a random seed, we still get random numbers that are different from each other, but they are reproducible. So whether we use a random seed or not, we still get random numbers. But with a random seed, we get a reproducible sequence of numbers. So keep this in mind: this is just like a little primer, and we will use this concept in a few moments. Slide 21 of 52, time stamp 19:26 So, for example, I mentioned before that we might get either “gray” or “overcast” if we randomly sample. Now, if we use a specific random seed like 42, we would always, for example, select “overcast”. I mean, it’s still a random selection, but we make it deterministic. In this case, given this prompt, the model will always select “overcast”. Slide 22 of 52, time stamp 19:49 If we use a different random seed, the model might select “gray”. Every time we sample, it will always select “gray” as the next token. So it’s still random sampling, but we are making it deterministic based on the random seed. Slide 23 of 52, time stamp 20:04 So, in watermarking, Claude watermarking is kind of like the idea that it sets a random seed. But this random seed, instead of being like a number that is fixed based on, I don’t know, someone writing down a fixed number, they’re using a secret key that is essentially like an API key, a secret key, and from that key, together with the four previous words, they derive this random seed essentially. But the idea is that if I go back one slide, it’s the same as here: there’s essentially a fixed random seed, and that random seed always selects the same next token. Okay, so instead of using random seed 99 here, for example, they have a secret key and also use information about the previous tokens to derive this random seed. But more on that later. Slide 24 of 52, time stamp 21:06 So the idea is that watermarking makes the text generation more deterministic in certain positions. So, for example, if we have these plausible texts on the left side. So if I have a text that says, > The weather today is cold and I may either pick “overcast” or “gray”. And the next sentence could be, > and then “light” or “gentle” They’re both interchangeable again. > And then breeze is “moving” or “blowing” through the trees, and the streets seem “quiet” or “still”. Which means basically I could say either “quiet” or “still”. So there are certain positions in the text where we have token choices where they are almost equally likely, like we have seen before. So that means if we are, this is without watermarking, if we are running the prompt, or given the prompt through the LLM, we might sometimes get this answer here, sometimes this answer, and so forth. And based on the number of positions, we might have 128 possible answers here. And of course, the longer the text, the more positions we have where we can have terms interchangeably, the more combinations, or the more output texts, there are. So, for example, again, one possible output text could be > The weather today is cold and overcast. A light breeze is moving through the trees, and the streets seem quiet. I think I’ll stay home and read a book with a cup of tea. So that is one possible text. Another possible text is > The weather today is cold and gray. A gentle breeze is blowing through the trees, and the streets seem still. I think I’ll stay inside and read a novel with a mug of tea. By the way, it’s also actually raining outside. I don’t know how good this microphone is, but it’s kind of a very fitting context here. But yeah, the bottom line is that you can see there are two very reasonable texts here being generated, and there are more combinations. So they are all reasonable. There isn’t one that is necessarily better than the other. They’re just, you know, slight variations. And if we don’t use watermarking, we might get either one, or it’s just random, right? Because of the random sampling, we might get one or the other. Slide 25 of 52, time stamp 23:31 Now, if we fix the random seed, as I mentioned before, for example, if the random seed is 99, we might always get this text here. So, using a random seed, we can kind of fix which answer we get, because then the random sampling is still random, but it’s deterministic in the sense that it’s reproducible. It’s always going to be the same then. Okay, so that is still without watermarking, now with a random seed. Slide 26 of 52, time stamp 23:59 And the watermarking is essentially doing the same thing. Now, instead of just using a simple random seed, they have a so-called random key, where this random key is involved in selecting the text, essentially. But what we can already say is that, in the Claude blog post, they say the watermarking shouldn’t make the text worse. If we look at this mechanism, yeah, it makes sense why it would not make the text worse. By the way, I’m not defending watermarks here. I’m just trying to explain. So please don’t kill the messenger here. But what I’m trying to say is that the watermarking is nothing else for the end user than fixing a random seed and making this sampling kind of deterministic, if that makes sense. Slide 27 of 52, time stamp 24:47 Okay, so the summary so far is without watermarking. We often sample without a random seed because I know most people don’t even use one. I honestly don’t think you can necessarily do it with the Claude and OpenAI APIs. I know you can do it in Ollama, but I also always had some problems with that because I used Ollama in one of my books for the bonus material to generate some texts. I was fixing the random seed, but it still wasn’t always deterministic, and so forth. So it’s tricky. Your mileage may also vary, depending on the software version and so forth. Anyways, so without watermarking, we have this random sampling. With watermarking on the right-hand side, we still have the random sampling. But in addition to just a random sampling being fully random, we have this watermarking key. And this watermarking key is passed to the random seed generator to set a specific random seed, making this deterministic. But it’s essentially very similar, and like I mentioned, there’s a lot of benefit in terms of understanding things from scratch. And now we know essentially where this watermark is applied to. So this is essentially applied to the sampling. It’s not applied inside the LLM, which is actually cool knowledge. So they don’t need to train a new LLM for that. They can just use an existing LLM, and they just apply it at this sampling stage. They don’t have to retrain anything or anything like that. So yeah, that is actually interesting, right? Slide 28 of 52, time stamp 26:21 But we are not quite done yet. I would also like to talk about how we can understand or see whether text is watermarked. So detecting the watermark is only possible if we have access to the key. So, for example, if we have these different texts, and essentially, after the text was generated, you find some random text on the internet (for example, you find this text number four here on the internet somewhere), you want to know: is this watermarked? Well, it’s impossible to know because, in order to know, you would need the watermarking key. You need this scoring function, and then you have to score basically the text with a scoring function. And then the idea is that if the score is above a certain threshold, then the text is watermarked. Otherwise, it’s not watermarked. But as the end user, we can’t do this because we don’t have this key. So the key is not available to us. Only Anthropic will have the key. However, in this blog post, they mentioned that they are providing it, of course, or they’re going to develop an API for that that they will make available. I don’t know. Honestly, I’m not affiliated. I don’t know the details. I was just reading this in this blog post. That’s all I know. So that API might as well be private for some companies, like, let’s say, X or Substack Notes, when they want to label AI-generated posts. They may make it public for end users to use. Who knows? We will have to wait on that. But yeah, so the bottom line here is that watermark detection is only possible if we have this watermarking key or, of course, the API that they are going to develop. Slide 29 of 52, time stamp 28:03 Now, removing the watermark is interesting. So now that we know how the watermarking works, we also know the shortcomings. I mean, this is really highly dependent on specific tokens in certain positions. So, for example, in this given text, if these colored words or tokens are the watermarking positions, we know that we could remove the watermark by editing this, right? If we change all the words at these positions, we would be 100% able to defeat this watermark. Now, the problem, though, is that we don’t know, right? Slide 30 of 52, time stamp 28:41 So we don’t know where these words are because we haven’t generated the watermark. So we don’t know which positions to look at. So the practical scenario here is that we could just randomly edit the text. So we would randomly change a few words and hope that we change enough positions to edit the watermark. So that would be one way to remove it. And since we also don’t know which are the highest-scoring ones, because that would require us to have access to the LLM and rerun the prompt through the LLM to find out which words are the highest-scoring, we can kind of only guess. So for example, we might say, oh, we replace “overcast” with “cloudy” because we don’t know that “gray” was high-scoring, you know? So in this case, it might be intuitive to say “gray”, but there might be cases where it’s not so intuitive. So what I’m trying to illustrate here is just some general text editing where we are modifying positions, but we are still kind of guessing what a watermark position is. So since we don’t know, we added just a few words here and there. And if we added enough words, that would also defeat the watermark. Slide 31 of 52, time stamp 29:59 So yeah, that was the watermarking in a nutshell. I mentioned that there is a scoring function to find out whether something is watermarked. And I want to do it as a bonus here. It’s already a long video, but as a bonus here, I wanted to briefly also explain how this scoring function works because that is also interesting information. It’s a bit complicated. It’s not essential to understand how the scoring function works. But the reason why they do it the certain way they do is to make the detection cheaper. Because otherwise, if I go back one slide or two slides, if you wanted to check if something is watermarked, if even they wanted to check, they would have to rerun the prompt to get these scores and then apply this watermarking random seed to get this text and then compare. And that would be very expensive because then essentially every text you want to compare, you would have to rerun the LLM. You have to know which LLM, and that would be really unfeasible because you often also don’t even know the prompt, right? So yeah, so they have like a trick that they use to, yeah, I would say, modify the sampling so that you don’t use or don’t need the LLM later on for the scoring stage. And in the blog post, they mentioned that they derived this method from a paper. It was a Nature paper, and this method is called SynthID-Text. So that was like a paper that came out maybe one or two years ago. It was by Google, and they use a similar technique they call Claude watermarking. I don’t know, sorry, I don’t know if they use exactly that technique, but that’s the one they mentioned. Slide 32 of 52, time stamp 31:39 So how does it work? So before we looked at the slides, we looked at the regular, let’s say, overview here, where we have some text. We put it through the LLM. We get this logit distribution and then we sample from the distribution and get the output token. And here, during the sampling, we use the watermarking key and the random seed generator. So this is still correct. This is still what’s going on, but there is a bit more, I guess, nuance to how this token is sampled. So they’re not just using, let’s say, NumPy’s random choice. They’re using something a bit more sophisticated here. Slide 33 of 52, time stamp 32:16 So assume, again, our context is “the weather today is cold,” and we want to generate the next token. So, for example: “gray”, “overcast”, “gloomy”, “cloudy”. “Gray” is 50% probability, “overcast” is 30, “gloomy” is 15. Let’s say “cloudy” is 0.05 and the rest is, let’s say, 0. Here it looks, of course, a bit different. Let’s say that’s “gray” and “overcast”. I’m just reusing this figure. But now imagine these are the most likely ones, like “gray” and “overcast”, and everything else is just very small, except “gloomy” and “cloudy,” maybe. So essentially, think about just a very small vocabulary for this example of four words instead of all these 50 words here, just to make it even simpler. Now, as I mentioned before, we could use ‘sNumPy’s random choice with these probabilities to sample the next token. Slide 34 of 52, time stamp 33:13 And we could use the watermarking key with this random seed generator to make it deterministic and get the certain watermark that we want. But as I mentioned before, this would be very expensive. Not the sampling itself. That doesn’t matter. This is pretty cheap. But the detection later on would be very expensive if we are trying to check random text on the internet. Slide 35 of 52, time stamp 33:35 So instead, what they use, they also use it during the sampling, during the generation, so that it can be reused later during detection. What they use is called tournament sampling. So this is instead of using something like random choice, they use a concept called tournament sampling. And so how does that work? It might look a bit complicated, but it looks really more complicated than it really is, to be honest. So you might have to, I guess, stop the video at some point and just sit with the figure a bit. But I think it is actually simpler than it looks like. It’s like once you get the hang of it, it’s pretty straightforward. But let me try to explain here. So what we have is we have still this context, and then we have these probable or plausible next tokens with these different probabilities. Now they have something they call random watermarking functions. Slide 36 of 52, time stamp 34:35 Here we have three watermarking functions, G1, G2, and G3. In reality, they might have 30, 50, or even more. Here I’m just using three because that is simpler on this slide. It’s just smaller, you know, like it fits better on the slide. Now, if we look at this word “gray”, this might give us a signature 101. With that, I mean, if we use this watermarking key to generate this random seed, and we have three functions, G1, G2, G3. If I put the word “gray”, what I’m skipping here is that usually you put the word “gray” together with the four or three previous words from the context. So it’s “cold” and “gray”. If I put that into G1 together with this watermarking key, I get the value one. Why? Well, that’s just how this function works. It’s like a random function. The random function either returns zero or one. In this case, with this random key and this token, it returns one. With the same key, but a different function, you get the value zero. And then here you get a one again. So if we have more functions (of course, 30 functions), this will be a very long string of ones and zeros. Slide 37 of 52, time stamp 36:00 It’s basically like a bit string, like if you have bits of zeros and ones. Okay. So this is for “gray”. So we get the signature 101 through using these watermarking functions. Now we can do the same thing for all the other ones. So we can do it for “gray”. We can do it for “overcast”, “gloomy”, and “cloudy”. So each one has a different signature here. So, for example, “overcast” is zero, one, zero. “Gloomy” has zero, zero, one. “Cloudy” has one, zero, zero. Okay. So we have these bits here now. The next step is a so-called tournament sampling where we just pair them. Slide 38 of 52, time stamp 36:39 Like, you know, like a soccer tournament, the knockout (KO) stages, or like the playoffs in American football, you always have two teams playing against each other. And that’s kind of like the same idea. We have a pair of tokens, and they’re playing against each other, essentially. And the scores, they come from these functions here. So we start with the first function in the first round. So we have “cloudy” and “gray”. So we look up here: “gray” is a one and “cloudy” is a one. Okay. So one and one. “Overcast” and “gray”. So “overcast” is zero, “gray” is one. So we have zero, one. “Gloomy” and “overcast”. So here we have “gloomy” zero, “overcast” zero. So zero, zero. And then we have “gray” and “gray” again, because we are running out. So we don’t have enough of the others. So we have one duplicate. So this is chosen randomly. And so you have one and one here. Now we look at the results. So this is a tie. In the case of a tie, we also select randomly using, you know, the random seed and the watermarking key. So here, “cloudy” survives. And from this one, G1 is, according to G1, “gray” is the winner because it has the one. So “gray” survives. And then here, “overcast” and “ gray “ are a tie, randomly selected, and “gray” also randomly selected. So we have now “cloudy” and “gray” and “overcast” and “gray”. And we play the next round in this tournament. So in this next round, we use G2. So according to G2, “cloudy” has a zero here. “Gray” also has zero. “Overcast” has one. And “gray” also has zero, sorry. And so, the next stage of the tournament again. Slide 39 of 52, time stamp 38:24 So we have a tie. We randomly select “gray”. And here we have “overcast” as the winner. And so we have “gray” versus “overcast” in the final. And then we look again at the scores. So “gray” has a one. “Overcast” is a zero. So “gray” is the winner. And that’s how the token “gray” is sampled. What is the watermarking key doing here? So the watermarking key, if I go back a few slides, is selected for generating these scores using these random watermarking functions. So the watermarking key determines essentially what values we get at these stages. So the watermarking key is still very important. Otherwise, these signatures would look different. Slide 40 of 52, time stamp 39:11 So we now have sampled the next token. And that’s just how this modified sampling procedure works. We could have used NumPy’s `random.choice`. But the shortcoming of that is that if we want to score random text on the internet, we would have to rerun the LLM. With this technique, we don’t. I will show you in a moment. So this technique sounds like really weird and cumbersome, but it has the advantage that we can now score random text more easily without having to rerun the LLM. So it’s essentially just to make the detection easier and cheaper. Slide 41 of 52, time stamp 39:43 Slide 42 of 52, time stamp 39:48 So, for example, if we have a new text. So I’m just using the same text here, but let’s assume it’s new text. So this is after the sampling, when we are scoring. And let’s say we are discovering this text on the internet. And the text is the weather today is cold and “gray”, and we want to know if this is LLM-generated or not. So we would, or Claude/Anthropic would, have the watermarking key and these functions: G1, G2, and G3. And it would put this text through these functions. For the one position here for “gray”, we would get 101, similar to what we got during the generation process. So this is the same as before. And this has, if we add up these bits, two bits, right? One and one here. So it has two bits of information, let’s say, for simplicity. This is just a really simple illustration. But let’s assume we get a score of two here for the “gray” in this position. If we had a different word here, “overcast,” in this position, we would get one if we get “gloomy,” like we also have one, and “cloudy” one. So I’m just summing over each row here, right? So that’s just like a score we would get at each position. And here I’m only looking at the last position. If I would do this at other positions, I would get a different score at different positions. So, for example, let’s assume at the first position I get a two. Here I get a two. For “today”, I get a three. For “is”, I get a two. “Cold”, two. And “gray”, three. So here I’m applying these watermarking functions as I’ve shown on the previous slide. Slide 43 of 52, time stamp 41:33 And I’m just adding up these numbers across the three functions. And the watermarking functions are very cheap. So you can just quickly run them on the whole text and get these scores. And then based on that, I can compute the average bits. So if I just average over all these values here, let’s say I get 2.23. Slide 44 of 52, time stamp 41:55 Now, if I have slightly different text, so here I swapped “today” with “now” and “gray” with “overcast”. These now get a score of one and one. And if I average over this whole string, then I get a 1.71. And so for that, I don’t need an LLM. All I need is the watermarking key, the random seed generator, and these functions, G1, G2, and G3. And that’s all I need. I don’t need the LLM. And I can get this score here. And what they do is apply a threshold. Slide 45 of 52, time stamp 42:27 So, for example, I mean, they don’t use this exact threshold. But for example, we can say if the score is greater than two, then the text is watermarked. If the score is smaller than two, it’s not watermarked. So here, if the score is greater than two, it’s a yes. So yes, this is watermarked. In this case, 1.71 is not greater than two. So this text is not watermarked. Okay. So that’s just the way we can then detect whether random text on the internet is watermarked or not. It’s essentially just applying these watermarking functions and then averaging over the scores and applying a threshold. Okay. Slide 46 of 52, time stamp 43:13 So again, the tournament sampling is mainly to make detection easier and cheaper. We could also use something like NumPy’s random choice with a random seed or to make the sampling deterministic. But then again, it would be hard to score any text on the internet. Slide 47 of 52, time stamp 43:30 So yeah, the summary is still the same, though. The thing that is different between no watermarking and watermarking is that we are controlling this sampling here with the watermarking key. And inside that, we have this tournament sampling. And yeah, as I mentioned before, detecting the watermarks requires the secret key and the watermarking functions G1 to Gn. Slide 48 of 52, time stamp 43:54 And again, removing the watermark, because I think that’s maybe interesting to some people, would ideally involve editing all the positions here. But since we don’t know which positions are watermarked and internally, they choose the positions so that they have equally likely tokens at those positions. And there might be positions where that’s not true. So here, for example, for “trees”, we might not even have an alternative word that is high scoring so they don’t watermark that position. So they only do the watermarking at certain positions essentially. Since we don’t know which positions to kind of defeat or remove the watermark, we would... Slide 49 of 52, time stamp 44:29 ...have to edit several places in the text. So what I think that means for the future of AI-generated text is that this actually... Slide 50 of 52, time stamp 44:36 …might result in worse AI-generated text. So I think if there’s a person who likes to use AI-generated text everywhere on the internet, let’s say there’s a news website that likes to use AI-generated text to write the news, I don’t think watermarking will necessarily stop them from doing that. They will probably still want to generate AI-generated text because that’s part of their workflow. So I think my guess is that they’ll use another model. Slide 51 of 52, time stamp 45:08 They’ll just use a second model to edit the text to get the so-called edited AI-generated text. So it’s complicating the pipeline. Instead of getting the text directly from Claude, it’s now using Claude to generate AI-generated text, passing it through a local model, and then having edited AI-generated text, which is likely not watermarked anymore. So why a local model? I just think a local model because I think all the providers- the proprietary LLMs, not only Claude, but also Google— I mean, Google wrote this paper, right? So I’m thinking that they are also watermarking Gemini text. And I think OpenAI is probably already doing it or will do so as well. I mean, I’m just speculating, but I’m imagining everyone will probably do something like that because there’s like an EU regulation that requires that. And that’s, according to the blog post, apparently why Claude is doing it. Yeah, so I’m thinking local models may not, at least not yet, implement this watermarking. So I think people will just use a local model and then generate edited AI-generated text. And my guess is it will be slightly worse than the original text because for the local model, you might now be using a smaller model. So, I mean, you could also technically just use the local model directly to generate text. But in my view, editing text is simpler than generating text. So for the generation of the text, you might use a very expensive high-end, I don’t know, like the highest, most expensive Claude model for complicated text. And then you use a cheaper local model to make these surgical edits, essentially. That’s probably what’s going to happen. And why worse? So if we look back at this graphic where we just added random positions, you might be just changing words for the sake of changing them. And then it risks making the text worse. So you might still have generated text, but it’s kind of like it’s edited awkwardly. Slide 52 of 52, time stamp 47:20 But anyway, so my goal here was to explain how the watermarking works and not, let’s say, the worldwide ramifications of that. But I hope this kind of behind-the-scenes, under-the-hood look is useful. The watermarking is not as complicated as it might seem, but I think it was still 52 slides, so it was also not super trivial. So I hope you found this little lecture useful. And yeah, until next time, see you then. PS: If you like more explainers in this style, I don’t post videos to YouTube regularly, but I have accumulated over 300 videos over the years, which you can find on my YouTube channel here . I also have a YouTube version if you prefer using the YouTube player And here is a link to the slides

0 views
Giles's blog 1 weeks ago

Use the built-in GELU, don't roll your own!

Unsurprisingly, PyTorch's own built-in GELU function is faster than the hand-rolled one I've been using to date. But I was surprised at how much faster using it made things when training my models. I discovered this accidentally just now while working on something unrelated, but am logging the details here for anyone else that might find it useful. The headline numbers: the same code, training the same model on the same data, ran at about: That's a 20% increase in throughput for both of the built-in versions -- definitely nothing to be sneezed at. And what is particularly interesting is that there aren't that many GELUs going on -- it's a GPT-2 small-style model, with 12 layers. So that's 12 GELUs handling tensors shaped , which is for my training setup. Given that the rest of the model is doing all of the normal full attention stuff for GPT-2, it's really surprising that the GELUs alone must have been taking up so much of the time. The throughput numbers mean that we must have been spending about 17% of our time on the extra overhead from the hand-rolled version, so that sets a lower bound for how much time the GELUs were taking up. More info below the fold. Back when I was doing the "interventions" part of my LLM from scratch series , training dozens of GPT-2 small-sized models in the cloud and on my local machines, to keep things simple I used the original model code from Raschka's book. That happens to have its own implementation of the GELU function -- you can see my copy here . I'm not that sure why the hand-rolled version is in there -- he covers the maths, but the specific implementation isn't explained in that much depth, and it seems rather like boilerplate, just a "type this in and use it" kind of thing. By contrast, for example, while he does explain the maths behind cross-entropy loss in similar detail, we use the built-in function for it rather than coding it up ourselves. When I switched to using JAX for my own from-scratch implementation , I decided to not bother porting the boilerplate, and just used JAX's own built-in version . I was revisiting the PyTorch code -- I'm in the process of extending it with mixture-of-experts support, about which more in a later post -- and decided to switch from the hand-written GELU to the PyTorch one just to tidy things up a bit. I noticed something interesting -- my new MoE code suddenly seemed to speed up. Was that a mirage? Or had I discovered part -- or even all -- of the reason why the JAX code was so much faster than the PyTorch code? With PyTorch, I was typically getting training speeds of about 21,000 tokens per second, while in JAX I was getting 24,000 tps or so. I'd been chalking that up to JAX's JIT compilation, but could it have been just a result of a random implementation choice I'd made? I did three partial test training runs, letting each one run for 20 minutes to allow the training speed to settle down from any startup overhead. Firstly, with the old hand-coded GELU: So it was getting 20,920 on average over those 257 global steps. That speed was in line with the original run of the configuration I was using. Next, I introduced the built-in PyTorch GELU with no arguments: That does the full calculations for GELU, rather than using the -based approximation that the hand-rolled code did. After 20 minutes, it looked like this: So this time we were getting 25,134 tokens per second -- 20% faster! By default, PyTorch's GELU uses an exact calculation of the function -- the hand-written code from the book uses an approximation using . Luckily, you can get that same approximation from PyTorch: So, training with that for 20 minutes: 25,142 tokens per second -- basically the same as the non-approximate version. So: switching to the built-in GELU made my PyTorch code run 20% faster, at about 25,000 tps rather than 21,000. My JAX code, which used JAX's built-in GELU, ran at around 24,000 tps. I'd actually found that rather surprising, because in JAX I was training in full-fat 32-bit floating point, while in PyTorch I was using Automatic Mixed Precision (AMP) -- a special mode that allows it to use 16-bit calculations where it won't hurt the model much. I'd found that AMP gave PyTorch a huge speedup -- from 15,402 tps to 19,797 on one test. So JAX without AMP being so much faster than PyTorch with AMP was a bit of a surprise. Its JIT is pretty amazing, but I didn't expect it to be that much faster. Now I think that we have at least part of an explanation. I was using JAX's built-in GELU (interestingly, with its default parameters, which means that it used the approximation), but the PyTorch code was using the hand-rolled one, and that unduly penalised it and erased some of the gains it got from AMP. If I really wanted to dig into this, I suppose I might try JAX with a hand-rolled GELU to see what happened. My guess is that because of its JIT, it might actually handle it better -- the whole hand-rolled thing could be compiled into one thing on the GPU. Perhaps it would also be interesting to try the non-AMP PyTorch code with the built-in GELU. But I doubt that would really be the best use of my time (and my electricity bill), so I'll leave it here. On the other hand, I do intend to have a look at in the future, to see what kind of speedup I can get from it. And it might be able to compile and fuse together the hand-rolled GELU -- so that would be an interesting thing to experiment with in that post: does the built-in GELU advantage disappear if we're compiling? But anyway, for now, lesson learned: use built-in PyTorch modules when you can. It's a pretty obvious one ;-) [Update] On X, Sebastian Raschka noted that he used the approximate version of GELU in his code so that the models were compatible with the OpenAI weights -- they were trained with that version, so they may behave slightly differently if you use the "pure" version. That's a great point, and so I've updated my own copy of the code to use . 21,000 tokens per second using the hand-rolled GELU from Sebastian Raschka 's book " Build a Large Language Model (from Scratch) ". 25,000 tokens per second using PyTorch's built-in GELU with no arguments. 25,000 tokens per second using the built-in GELU with , which uses the same maths as Raschka's version under the hood.

0 views
Pete Warden 2 weeks ago

Why I ported Moonshine to Javascript

One of the most common requests I’ve heard from developers is an in-browser version of Moonshine that can run on a web page. In theory this should be straightforward – we already built MoonshineJS for the previous generation of models, and the core library is written in portable C++, so emscripten can compile it into WASM. There have even been some interesting community porting projects but I held off on official support until I had time to do it justice. I knew that porting the C++ core was just the beginning. Building something that would be straightforward for web developers to use required a lot more: After a lot of work, I finally have a version ready for feedback. The easiest way to try it is on the new moonshine.ai home page, where you can now see everything from a minimal transcription example to a full-blown Granola-style meeting note taker . As an open-source project, all the code for these is available and the simple examples include code snippets in-line too. Here’s one that shows how to run speech to text on a web page, to give you a flavor of the API: You may still be asking yourself why I made supporting Javascript in the browser such a priority? A lot of “X ported to WASM” stories end up being Hacker News bait without having any practical uses. The evidence that drove me was: I’m excited to get feedback on how to improve the initial version, and I’m looking forward to hearing about what people build with it, so please come by our Discord channel if you’d like to join our community. High-level APIs that were both idiomatic for browser Javascript and consistent with the other Moonshine language bindings. Infrastructure for testing from units to full web pages. Integration with the existing CI and deployment process. Examples that were interactive and showed the key capabilities of the library, with interactive inline code. Larger applications that demonstrated and tested how the framework runs in real-world conditions. Improved support for in-memory models and data files. This was involved a lot of changes to the core library, because while there had always been some methods that took memory buffers, coverage was patchy compared to loading from files. Clear developer demand . It came up frequently as a wishlist item when talking to users. Javascript’s dominance . Python rules machine learning, but JS is the most common language for applications, web and server-side. Advantages over alternatives . Voice interfaces are clearly only going to grow in importance over the next few years, but browser APIs are neglected and server-based alternatives are slow and costly compared to our on-client framework. Obvious applications . Dictation and meeting note taking are popular use cases for speech technology already, and talking with AI bots is becoming a lot more common too. Technical alignment . Deep in my bones I know that voice interfaces want to run on the client. The current status quo of streaming audio data to a server just to get text and intent back only exists because models used to be too large to run on consumer hardware. Today even household appliances have enough compute horsepower for local voice agents . It offends my engineering sensibilities to see old approaches linger on purely out of inertia. Speech wants to be free to use and local, just like all our other input devices like keyboards, mice, touchscreens, and cameras. Porting makes that possible on the web. Options . These days a lot of us have to frequently switch between languages and operating systems, and the power of AI coding assistants only increases the pressure to rapidly port applications. A library dependency is a big commitment, and knowing that it will be available anywhere you’re likely to run in the future makes the risk of betting on a framework much lower, even if you don’t need the option in the end.

0 views
Farid Zakaria 3 weeks ago

nixpkgs-multiverse: every version that ever existed

Enter the Nixpkgs multiverse. All the versions that ever existed, all in one place. I bumped the release for my NixOS configuration to refresh many of my packages and found that a package I depended on at a particular version is no longer available. The package was “version bumped forward” in a way that broke some of my tooling. It’s late and I don’t want to fix it, so I just add another input pinned to the commit that had the version I want. This works, but it is miserable in a way that compounds. The need for the most recent package is so common that I had keep an overlay that would inject as a package set for me to easily pull from. If I have a need for a particular version of a package and it’s not present in my current , I am left searching for the commit and pinning it. 1 Every pin is a whole extra in the file. Flake inputs are fetched eagerly even if not used. A flake with three inputs whose output references only the first are all materialised. Nix lets us easily create a closure that reproduces a specific version of a package, but Nixpkgs makes it hard to hold one package still while everything else moves. Each Nixpkgs input to a flake is a distinct universe. If we can have multiple Nixpkgs as input to achieve fetching a particular package, why not have every version that ever existed always available ? 🤯 nixpkgs-multiverse is one flake input that gives you all of them at once. We can query the flake for all the versions of a package that ever existed in Nixpkgs . If we want a specific complete revision of Nixpkgs we can use the function. That is access to all the versions of all the packages that ever existed in Nixpkgs. You can mix them all together in one shell, one package or a build environment. How is it possible to have multiple Python versions? That is the whole point of Nix itself. Every package immaculately describes its dependencies using a hash via the intensional model . 2 Nixpkgs already supports multiple versions of a package in a single revision (i.e. , , ) as separate attributes. We took this to its logical conclusion of making them all available easily. Our deliberately has no inputs: . Inputs are fetched eagerly, and we have 1,393 of them. We need to fetch them lazily, only when something actually references a revision. To do this, we fetch revisions with , pinned by , only when needed. Two files do all the work: and . is one ordered array of every revision from Nixpkgs , 1,393 as of this writing, from 2017 to 2026. 3 We limit our commits to those that were actually built and cached by Hydra, so we only include commits that were either a release or a channel bump. How do we know which revisions to pick for the ones? We rely on the nix-releases S3 bucket to tell us which commits actually became published builds. The S3 bucket uses the commit hash as the directory name, so we can list the bucket and get a complete list of all revisions that were actually built. is the map from (attribute, version) to a revision: That integer is an offset into . It is the most recent revision that shipped that version. At this many revisions, it turns out that how you encode the data matters a lot. My first encoding stored every revision a version appeared in. Although it was simple, it was a disaster in terms of size for these JSON files. As you might expect, most versions of most packages are unchanged across many revisions. The size of our file was growing linearly with the number of revisions. By storing only the newest revision that shipped a version, we can keep the file small and still answer the question “which revision had this version”. Here is how it actually grows as revisions get indexed: 5.18 MB covering 1,393 revisions and 289,521 distinct (attribute, version) pairs. The key design rule for our flake: Cost is per revision touched , not per package. If we were to add revisions as inputs, evaluating our flake would explode. Each flake in our measurement below has N inputs and an output that references only the first one ; the timing is how long before that output evaluates. 4 Five pins that are not used cost 26 seconds before the output evaluates. Each input costs about 5 seconds, and the input is fetched and materialised even if never used. In contrast, the green line is with 1,393 revisions available , which is a flat 0.20s to parse the JSON. 🤩 Revisions are memoised, so pulling 3 packages out of one revision costs the same as pulling one. That concept that the can hold many graphs of the same package is core to understanding Nix. The popularity and rise of flakes made it even more apparent that we can mix multiple revisions of Nixpkgs together. The thing I keep coming back to is that Nixpkgs history already is the multiverse. Every version that ever existed is already built, already cached, already reachable. It was just addressed by commit hash instead of by version number, which is exactly backwards from how anyone thinks about it. The whole project is 5 MB of JSON and about 200 lines of Nix. It does not build anything, mirror anything, or host anything. It is a phone book. Thankfully sites like nixhub.io or lazamar’s search make this a little easier.  ↩ The hash is a unique identifier for the exact set of inputs that were used to build it. If you change any input, the hash changes and you get a new package.  ↩ A NixOS release is not special; it is a commit that happens to carry a label.  ↩ Everything is against a local clone, so there is no network latency.  ↩ Thankfully sites like nixhub.io or lazamar’s search make this a little easier.  ↩ The hash is a unique identifier for the exact set of inputs that were used to build it. If you change any input, the hash changes and you get a new package.  ↩ A NixOS release is not special; it is a commit that happens to carry a label.  ↩ Everything is against a local clone, so there is no network latency.  ↩

0 views
James O'Claire 3 weeks ago

How to Install Python3.14 from source on Ubuntu 26.04

I’ve always been a big fan of installing Python from source. It keeps you close to your environments and understanding exactly where the code comes from and goes to. I feel like in the past it was more difficult, but year after year it seems like it’s getting easier. This year for Ubuntu 26.04 there is a bit of a new step added, but ultimately it is feeling effortless. The whole process takes about 10 or so minutes, and is good for people of all levels to understand. You may need a C compiler if you haven’t yet downloaded: There is a new way to get the build-deps for Python on Ubuntu, so you’ll need to edit and add to the Types line. Now that has been added you can get build-deps for Python: I generally install all additionals since many are quite important for python. Some are more specialized and you may not use, while others like libcurses are basically required if you expect to ever need to interact with an interpreter (eg troubleshooting a production environment). Others you may not need until some future data, say when you install a package that expects a certain kind of data compression. Overall, my advice is just install all unless you know more or have a specialized use case for streamlined environment. Download the latest version: Choose Gzipped Tarball https://www.python.org/ Python is now installed. Last important step is to configure the system links: x Finally you can test: you will now have a directory like Python3.12.xxx Move python installation directory to /opt/ to keep home clean into the directory and run each command Check you have all packages installed from above Builds Python, this step takes some time. altinstall skips creating the python link and the manual pages links, install will hide the system binaries and manual pages. This means that we will leave the system python installation untouched

0 views
Evan Schwartz 3 weeks ago

Notes from the AI Coding Transition

Like many other software engineers, my coding workflow has changed dramatically since the start of 2026. And like many others, I've felt some mix of awe, grief, frenetic productivity, atrophying skills, and understanding less while shipping more. In this moment where the field is undergoing this rapid shift, I've found it helpful to read others' takes on their processes, what they're doing to keep their brains engaged, and their genuinely mixed feelings. Before writing up my own thoughts, I went back through the relevant essays and blog posts from the last ~7 months to find the ones that resonated with me the most. Below are the posts that I especially liked and lines that stuck out from them, either because they gave me some idea about how I might want to use AI or just because they had a particularly incisive description of our field's situation. (Quotes are exact and the bold text is my added emphasis.) If you've read others that you thought were particularly on point, please send them my way! I didn’t ask for the role of a programmer to be reduced to that of a glorified TSA agent , reviewing code to make sure the AI didn’t smuggle something dangerous into production. If you would like to grieve, I invite you to grieve with me. We are the last of our kind, and those who follow us won’t understand our sorrow. Our craft, as we have practiced it, will end up like some blacksmith’s tool in an archeological dig, a curio for future generations. Even if AI agents produce code that could be easy to understand, the humans involved may have simply lost the plot and may not understand what the program is supposed to do, how their intentions were implemented, or how to possibly change it. Peter Naur reminded us some decades ago that a program is more than its source code. Rather a program is a theory that lives in the minds of the developer(s) capturing what the program does, how developer intentions are implemented, and how the program can be changed over time. Cognitive debt tends not to announce itself through failing builds or subtle bugs after deployment, but rather shows up through a silent loss of shared theory. As generative and agentic AI accelerate development, protecting that shared theory of what the software does and how it can change may matter more for long-term software health than any single metric of speed or output. the sense of psychological ennui leading into existential dread that many software developers are feeling Simon: All of the chess players and the Go players went through this a decade ago and they have come out stronger. The Shen-Tamkin study identified six distinct AI interaction patterns among developers. Three led to poor learning: full delegation, progressive reliance, and outsourcing debugging to AI. Three preserved learning even with full AI access: asking for explanations, posing conceptual questions, and writing code independently while using AI for clarification. The differentiator wasn’t whether developers used AI, it was whether they stayed cognitively engaged. metrics don’t capture what’s happening underneath. The mental fatigue of reviewing code you didn’t write all day. The boredom of babysitting an agent instead of solving problems . The slow, invisible erosion of the hard skills that made you good at this job in the first place. You stop holding the architecture in your head because the agent handles it. You stop thinking through edge cases because the tests pass. You stop wanting to dig deep because it’s easier to prompt and approve. There’s no spark in you anymore. Here is something that gets lost in all the excitement about AI productivity: most software engineers became engineers because they love writing code. Not managing code. Not reviewing code. Not supervising systems that produce code. Writing it. The act of thinking through a problem, designing a solution, and expressing it precisely in a language that makes a machine do exactly what you intended . That is what drew most of us to this profession. It is a creative act, a form of craftsmanship, and for many engineers, the most satisfying part of their day. this is different because it is not asking engineers to learn a new way of doing what they do. It is asking them to stop doing the thing that made them engineers in the first place and become something else entirely. a mid-level backend engineer is now expected to understand product strategy, review AI-generated frontend code they did not write, think about deployment infrastructure, consider security implications of code they cannot fully trace, and maintain a big-picture architectural awareness that used to be someone else’s job. That is not empowerment. That is scope creep without a corresponding increase in compensation, authority, or time . From my experience building and scaling teams in fintech and high-traffic platforms, I can tell you that role expansion without clear boundaries always leads to the same outcome: people try to do everything, nothing gets done with the depth it requires, and burnout follows. Now the only limit is your cognitive endurance. And most people do not know their cognitive limits until they have already blown past them. Set explicit boundaries around role scope. If you are asking engineers to take on product thinking, planning, and risk assessment in addition to their technical work, name it. Define it. Compensate for it. Do not let it happen silently and then wonder why your team is burned out. talk about what you are experiencing. The isolation of feeling like you are the only one struggling with this transition is one of the most damaging aspects of the current moment. You are not the only one. Computer programming is, fundamentally, about two things: I have a hard time imagining a future where knowing how to solve problems with computers and how to control the complexity of those solutions is less valuable than it is today, so I think it will continue to be a viable career even with the advent of AI tools. I try not to use LLMs to generate full solutions that I am going to need to support. Whenever I have Claude do something for me, I feel nothing about the results. It feels like something happens around me, not through me. the default output has no soul. It's correct. It's competent. It's fine . And "fine" is the enemy of everything I care about as a writer and an engineer. find it hard to believe that supervising a set of agents is going to lead to an optimal flow experience, because we are more passive, it doesn’t stretch our abilities in the same way, and it requires far less concentration. Will we find flow elsewhere? Solving problems and delivering value will always be rewarding, but I wonder if the optimal flow experience offered by programming has, for the most part, disappeared forever, and many of us will simply find less enjoyment at work. You realize you can no longer trust the codebase. Worse, you realize that the gazillions of unit, snapshot, and e2e tests you had your clankers write are equally untrustworthy. The only thing that's still a reliable measure of "does this work" is manually testing the product. Congrats, you fucked yourself (and your company). You let them run free, and they are merchants of complexity. They have seen many bad architectural decisions in their training data and throughout their RL training. You have told them to architect your application. Guess what the result is? An immense amount of complexity, an amalgam of terrible cargo cult "industry best practices", that you didn't rein in before it was too late. All of this compounds into an unrecoverable mess of complexity. The exact same mess you find in human-made enterprise codebases. Those arrive at that state because the pain is distributed over a massive amount of people. The individual suffering doesn't pass the threshold of "I need to fix this". The individual might not even have the means to fix things. And organizations have super high pain tolerance. But human-made enterprise codebases take years to get there. The organization slowly evolves along with the complexity in a demented kind of synergy and learns how to deal with it. With agents and a team of 2 humans, you can get to that complexity within weeks. And I would like to suggest that slowing the fuck down is the way to go. Give yourself time to think about what you're actually building and why. Give yourself an opportunity to say, fuck no, we don't need this. Set yourself limits on how much code you let the clanker generate per day, in line with your ability to actually review the code. When people say “taste,” what they actually mean is experience. Pattern recognition built up over years of doing the work. But calling it “taste” instead of “experience” does something subtle and harmful: it makes a learnable skill sound like a gift . Doing tasks manually naturally builds up the context required for the decisions involved later because you have time to process everything along the way and construct your mental model of the project's structure. This process requires more attention and context switching, along with way more decisions per hour. Making constant architectural, big-picture decisions while overseeing the work of a cracked junior dev is fundamentally harder than executing standard programming tasks yourself. Decision fatigue is, in my opinion, the next invisible friction point for developers. The problem is that as the coding agents get more reliable, I’m not reviewing every line of code that they write anymore, even for my production level stuff. But I’m not reviewing that code. And now I’ve got that feeling of guilt: if I haven’t reviewed the code, is it really responsible for me to use this in production? There’s an element of the normalization of deviance here—every time a model turns out to have written the right code without me monitoring it closely there’s a risk that I’ll trust it at the wrong moment in the future and get burned. When you stop fighting with hard problems directly, the mental models fade. You stop building intuition. You start pattern-matching on outputs instead of reasoning from first principles. And the worst part –> you don’t notice it happening. The code still ships. The PR still merges. Everything looks fine until the incident at 2am where you genuinely cannot reason about what the system is doing because you never really had to learn it. There’s a good analogy here from aviation. Pilots trained heavily on autopilot gradually lose the ability to fly manually and this isn’t theoretical, it’s contributed to real crashes. I think judgment is built from a specific loop: you form a view, you commit to it, you see what happens, and you update. That cycle, repeated enough times, is what builds calibration. The problem with AI is that it short-circuits the first step. You skip forming your own view and go straight to evaluating someone else’s. Do that enough and the muscle atrophies and again, not dramatically, just quietly. You become a better reviewer and a worse thinker. I did the software engineering equivalent of forwarding an email with “thoughts?” and then going to lunch . The job is the part where your fucking brain has to be in the room. You paste the issue into the machine before reading it. You accept the explanation before forming your own. You create a PR before even understanding what the problem you’re fixing is (!). You request a PR review before reading the diff. You merge because the checks passed and the reviewer approved it and the whole thing smells like progress. here’s the new hard rule I’m following after this “incident”: if I still can’t explain the change, I can’t ship it . No exceptions. many software engineers labor under a delusion that their job is to be excellent at their craft. Of course, wanting to be an excellent programmer is not a delusion; it is a completely legitimate value to hold, and a legitimate purpose to pursue. It’s just not what you’re paid to do at work. Your job , unfortunately, is producing shareholder value . This delusion has been punctured by the end of ZIRP , and again more recently by the rise of AI coding. Today, the ownership mindset defines the role. Although unintuitive, limiting the amount of work that runs in parallel is actually producing better outcomes and outputs. I believe the idea of WIP limits must be emphasised more strongly than before. moving from building features in parallel to building a single feature end-to-end faster. But for me, prolonged use becomes insidious. It's easy to become lazy and hand over thinking to the machine in looking for the next hit of cognitive offload when coding becomes even a smidge difficult. Why type your search and read half a short blog post to understand the problem when the same keystrokes give you the (possible) answer right there and then. When you ask a person to do something, you don’t expect them back in five minutes saying it’s done and ready for the next task. With an agent, that’s exactly what happens. Done. Next. Done. Next. There’s no breathing space. There’s always a next thing to think about. The work used to have a rhythm to it. You’d struggle, you’d get stuck, you’d finally figure it out, and there was this moment of joy when it clicked. Hours in the code, and then done. Figuring it out was the whole reward. That’s what AI can quietly take from me. Not the joy itself, but the sense that the thing was mine, which is where the joy was coming from all along. It hands me the finished thing, the finished thing works, and somewhere in there, I stop being the person who made it and become the person who approved it. AI didn’t take the joy out of coding, I gave it away. a quieter admission: the work isn’t teaching me much anymore, and it’s stopped being fun. That’s a description of becoming a manager . What AI did was give every engineer a small team of tireless, fast, occasionally-wrong direct reports. And with the team came the manager’s problem. The discomfort engineers are feeling right now isn’t an AI problem. It’s a delegation problem, and delegation is the oldest unsolved problem in our discipline. The good news: it’s not unsolved because nobody tried. Managers have been failing at it and slowly adapting for decades. What is in your control is small and it is everything: where you point your attention, what standard you hold, what you decide not to do, and whether you’re honest about which is which. The whole reason “there is too much” feels like drowning is that we keep trying to exert control over the size of the ocean. You can’t. You can only decide where to swim. I want to be able to explain what the system does without first having to ask a clanker to explain it to me. Present-day models tend to produce code that is too defensive, too complex, too local in its reasoning. They avoid strong invariants. They add fallbacks instead of making bad states impossible. They duplicate code, invent bad abstractions, and paper over unclear design with more machinery. If each iteration adds another small defense, the system slowly becomes less understandable while appearing more robust. we may no longer understand the whole system in the same way. We treat it, we monitor it, we stabilize it, but we do not necessarily comprehend it. Some domains will punish sloppiness and demand trust and responsibility, but a lot of software lives in a world where raw speed, quick experimentation, and vast coverage matter enormously. Better visualizations of changes or orchestration or agents will not restore our understanding. Either we need to find clever ways to jolt the human back into the loop and make the changes of the loops legible long term, or we need to find better ways to compose these ever more complex systems. In the old workflow, the creative process happened mostly in your mind. In the new process, you supervise the creative process that unfolds inside the AI’s internal machinations. Now, let’s put the historical novelist in the position of the software developer. She gets a call from her publisher saying they’ve found a way for her to bring four books to market each year instead of one book every two years. They’ve recruited a bunch of top-notch high school and college students who can each crank out five pages a day of competent writing for dirt cheap. The publisher wants the historical novels to maintain the original writer’s level of excellence, or to at least be close, so they’re retaining her services as an editor. The novelist’s job is now to edit the work of the students, each of whom has been carefully prompted to write pages that should, with a little work, be stitched together into coherent chapters. Anyone who has ever graded the work of high school and college kids knows that this is generally not rewarding work. If you’ve ever had to grade a hundred papers in a week, you know what a grind that is. The novelist, like the software engineer, is no longer deeply engaged in her work. Editing is not creating. You do not give yourself over to your imagination. You do not immerse your mind and feelings in the process of invention. Instead, you’re rooting out problems, trying to clean up clumsy wording and redundant descriptions instead. The flow state is gone. You are now a cog in a larger process that doesn’t really value your creativity or your need to exercise it. Worse still–and I have felt this personally after months of reviewing AI-generated code–your skills drop off sharply. When a new issue arises–a feature to be implemented, or a tricky bug to fix–the idea of wasting several hours on it feels insulting. Why should I dig through all that code when Claude can locate the bug in five minutes and start drafting a fix? But I think that creative people choosing to hand over their most imaginative, flow-state thinking to an army of bots will be a mistake in the long run. The feature gets delivered, but I do not really feel like I built it. Maybe this is just another evolution of our profession and in a few years it will feel completely normal. Or maybe one day we will realize that somewhere along the way we stopped programming and nobody really noticed. “I’m not sure I can do my daily job without Claude” The cost was never writing the code. The cost was owning it. A fix you cannot judge, in code nobody on your team understands, is not maintenance; it’s another spin of the roulette wheel . And when the bug comes back wearing a different hat, who do you escalate to? Your vibe-coded grid has no changelog, no support contract, and no team whose reputation depends on it. AI makes touching the code cheap; it does not make answering for it cheap. we are yet to see “mind blowing” software being churned out showing that it is still hard to build great software purely with agents. Coding using models can take you from 0 to 1 very fast. But what about 1 to 10, 10 to 100? In sufficiently large codebases, everyone operates with an incorrect theory of the program . Like many software tools, LLMs are a double-edged sword: they make it harder to construct a detailed mental theory of the software, but they allow you to build a partial theory quickly and they can help you leverage that partial theory more effectively. This is a complex tradeoff that I’m still thinking about. our field is evolving in an incredible and painful (but also joyful) direction if you control the ideas of your software, looking at the code itself is suboptimal and often pointless. large software projects have never been limited only by how quickly an individual can produce code. They are limited by how well people can coordinate their understanding of the system they are changing. The shared language of a software project is not English or Python but it is the common understanding of what its concepts mean, where the boundaries are, which invariants matter, who owns what, and why the system has the shape it does. Before agents, some of this shared understanding was maintained by friction....Some of it was the process by which your understanding became mine, and by which both of us discovered whether we still agreed about how the system worked. The most important skill in prompting is expertise in the domain you’re prompting for. A good illustration of this is Terence Tao’s conversation with ChatGPT about the recently-discovered counterexample to the Jacobian Conjecture. This is not the same ChatGPT I talk to! I couldn’t get to where Tao gets , even with unlimited tokens to burn. There’s a lot to learn about good prompting from Tao’s conversation. Here are a few observations: So why does software keep getting worse across the board? The bar for “user experience” has kept rising, but everything has become increasingly fragile. LLMs are useful for producing code that meets easily and objectively verifiable acceptance criteria which you provide explicitly I've found this simple instruction to vastly improve LLMs' output: " Never write READMEs, docstrings, or comments. I will write those myself later. And yes, I really mean this." Problem-solving using computers Learning to control complexity while solving these problems Write before you look. Before opening a tool, before asking the model, write down what you think. Not a design doc necessarily, just your current understanding of the problem, your instinct about the solution, where you think the tricky part is. Even a few sentences. This forces you to articulate your reasoning rather than pattern-match on someone else’s output. It’s also surprisingly useful as a diagnostic: if you can’t write anything, you probably don’t understand the problem well enough to evaluate any answer. Form a view before reading the suggestion. When reviewing AI-generated code or design, read it critically with your own opinion already in hand. What would you have done? Where does this differ? Why might the model have gone this direction and is it right? This sounds small but it’s the difference between passive consumption and active evaluation. One builds judgment, the other just builds familiarity with AI output. Separate ownership from authorship....You can own code you didn’t write. You cannot own code you refuse to understand. Those are different statements, and the gap between them is the whole job. Decide what you must understand deeply - then triage the rest without guilt. The discomfort is the job, not a bug in it. Acting on incomplete information, sitting with the unease of not-fully-knowing, and committing anyway - that is judgment. Managers don’t feel more certain than you; they’ve made peace with feeling uncertain and moving regardless.... Keep something you understand deeply.... Track what you’re learning, not just what you’re shipping.... Tao’s messages are very short and to-the-point. He doesn’t respond point-by-point to the model, just to the gist The model outputs are much more concise than when I try and talk to GPT-5.6 Sol about mathematics. By signalling expertise, Tao shunts the model into “talking-to-mathematicians” mode, not “explaining-to-amateurs” mode Tao pushes back when the model’s responses look wrong, but he doesn’t directly contradict; instead, he says things like “this looks more complex than I was hoping for” Tao makes several leaps and suggestions himself. He almost never takes the model’s advice about where to go next

0 views
Simon Willison 3 weeks ago

New release of LLM adds support for reasoning traces, OpenAI Responses, server-side tools, and smarter logging

I released LLM 0.32 this morning, the most significant new version of LLM since the initial launch of the project. The new version includes support for visible reasoning traces, server-side provider tools, redesigned content-addressable SQLite logs, new models, and new features enabled by the OpenAI Responses API. I also released a new version of the llm-anthropic plugin with substantial updates of its own. Running LLM against reasoning models now displays their reasoning traces to standard error, so you can see what they are "thinking" without that information being included in the standard output that you might pipe to another tool. Add to turn this off. LLM includes support out-of-the-box for the GPT-5.6 model family , and the new default model used with is now the inexpensive but capable GPT-5.6 Luna . LLM calls can now use server-side tools from various providers. OpenAI provide a code execution environment as a server-side tool; LLM can now run prompts that benefit from that like so: OpenAI also gets a WebSearch tool. The llm-anthropic plugin adds WebSearch , WebFetch , CodeExecution , and AnthropicMCP , which looks like this: That causes Anthropic to execute MCP calls against my new datasette-mcp plugin as part of a single request/response interaction with their API. The new llm openai endpoint command provides a tool for executing prompts against any OpenAI compatible endpoint as a one-liner. These aren't logged, which makes this a handy tool for running one-off prompts against anything that speaks the lingua franca of the LLM API world. Here's how I use that to run prompts against Gemma 4 12B running in my localhost LM Studio API, via (no LLM installation required) and mixing in the llm-tools-quickjs tool plugin for good measure: LLM's Python API previously required you to create a conversation and then send messages to it one at a time. This was an abstraction over the true nature of LLMs, where each request carries a complete history of the messages that came before it. That abstraction started to get in the way for some more advanced cases, so the new release introduces a parameter that can be used like this: LLM previously returned an iterable sequence of strings from each prompt. This worked great when models returned a string response, but failed to predict the weird shape that models would evolve towards. Today many models return a mix of reasoning text, output strings, tool calls, and even image attachments. With LLM 0.32 you can do this instead : Combine these features and we can finally provide a robust implementation of the semi-standard OpenAI chat completions API, which I've now released as the llm-chat-completions-server plugin: Now you can run prompts against LLM via that server, using the new command! The bigger challenge with that kind of API concerns logging. If we're going to support the pattern where the message sequence is appended to on every request, ideally we can avoid logging all of that duplicate JSON for every turn. The solution is the new content-addressable message store , modeled after Git. You can see the new schema for that in the documentation , but the and commands have both been upgraded to convert that format back into something that's easy to consume. There is a whole lot more in this release. The 0.32 release notes are pretty comprehensive, and the notes for 0.32rc2 , 0.32rc , 0.32a3 , 0.32a2 , and 0.32a0 should fill in any gaps. Existing LLM plugins should all continue to work, but plugins that provide extra models will need to be upgraded to 0.32 in order to participate fully in the new streaming events system. There's a guide to implementing plugins with Structured messages and streaming events in the documentation. I've updated some of my own plugins: Quite a few of the lower-level tools changes in this release were driven by the needs of Datasette Agent . When I started work on LLM, the term "agent" had such a vague definition that I refused to use it. In September 2025 I came around to the idea that " An LLM agent runs tools in a loop to achieve a goal " is well established enough now that I could stop avoiding the term entirely. Tool chains can now pause for human approval and resume from a stored message history - both needed by Datasette Agent. Looking at LLM today it's beginning to look very agent-shaped to me. There's something neat about having a CLI utility that can mix and match different tools from different sources with different models all as a one-liner, and that includes a Python library powerful enough to build systems like Datasette Agent and llm-coding-agent . Maybe the next version of LLM will bake the concept of an "agent" into the core library. I'm still trying to figure out what that would look like. You are only seeing the long-form articles from my blog. Subscribe to /atom/everything/ to get all of my posts, or take a look at my other subscription options . llm-anthropic 0.26 adds support for the Claude 5 family of models, plus , , , and server-side tools. llm-gemini and llm-openrouter and llm-mistral are nearly there, releases coming soon.

0 views
マリウス 4 weeks ago

The TEMU-fication of Software, Digital Goods & Services

Disclaimer: This is an opinion piece and most of it is speculation about a future that has not arrived (yet?), based on a few data points that have. As usual, summary at the end. A few years ago I would have laughed at anyone telling me that there is a serious market for ten-dollar drills, two-dollar dresses, and one-dollar pairs of shoes shipped from a warehouse on the other side of the planet. Today, however, that market exists and it has a name, and it is even publicly traded (sort of, through holdings). TEMU , Shein and a few others have built frankly mind-boggling businesses around the idea that if you make production cheap enough, fast enough, and just barely good enough to look right on a phone screen, an enormous part of the population will buy it, even when the product breaks within a week, when the materials it is made of contain worrying levels of toxic substances , and when the carbon footprint of one delivery exceeds that of an equivalent local purchase by orders of magnitude. The key to this sort of business model is not innovation, but instead the externalization and compression of cost. Somewhere upstream, people work seventy-five hours a week , in conditions most readers of this website would refuse to even visit, so that the rest of us can have a cheap plastic spatula at our doorstep within five business days. While the visible price collapses, the invisible costs get distributed onto landfills, lungs, and ultimately people that we will never meet. What follows is a hypothesis I cannot prove but have been turning over in my head for a while, as we are watching the same thing happen to software, books, music, (film-)scripts, and most of the digital goods and services we consume. The cheap labor in this case is not human, it is a Large Language Model ( LLM ), or what many people these days call “AI” , and the externalized cost is, among other things, quality , which requires craftsmanship to produce, and attention to perceive. And just like with physical goods, we will probably end up with a two-tier market, in which we have a large and massively profitable lower tier of generated slop , and a smaller, more expensive upper tier of work that is still recognizably human. I’d like to call this the TEMU-fication of software, digital goods and services , and describe what it might look like. For decades, the global fashion industry has relied on a workforce that has almost no leverage and no voice, and for which the economics work because someone, somewhere far away, will sew a t-shirt for less than the price of a coffee. Without that skewed arrangement, the entire fast fashion business model collapses. The garment in your hand is only cheap to you because it has been expensive to someone else , in ways that the price tag does not show. Modern Large Language Models occupy a similar position in the economy, with one important difference, which is that there is no human being in the sweatshop, only a stack of GPUs trained on a corpus of work that other human beings produced over the course of decades. The labor that has been compressed is historical and the model is a kind of compressed copy of the work of millions of programmers, writers, illustrators, and musicians, served back at near-zero marginal cost. Well, at least in theory, and only if the hyperscalers find a way to lower the cost per token, but that’s a different topic. However, the result is the same. A class of goods can suddenly be produced for an order of magnitude less than before. And, just like with TEMU , those goods turn out to be just barely good enough . The most direct manifestation of this so far is what is being called vibe coding . The term refers to the practice of describing what you want in natural language to an LLM , accepting whatever it produces, iterating over it with more refined descriptions of the basic idea and eventually shipping the result into production. Whether the developer actually understands what was generated is increasingly considered an implementation detail . And while the output is technically software, the question is what kind of software it is. A 2025 Veracode report found that approximately 45% of AI-generated code samples failed security tests and contained critical vulnerabilities from the OWASP Top 10 , and a multi-language, multi-model academic study that evaluated outputs from Claude , Gemini , Codestral , GPT-4o and Llama-3 across Python, Java, C++ and C, found that a substantial fraction of generated snippets were either non-compliant with basic secure coding standards or actively triggered classified weaknesses (buffer overflows, hard-coded credentials, SQL injection, cryptographic misuse, path traversal, you name it). Even more concerning is a peer-reviewed 2025 paper from IEEE-ISTAS that documents a 37.6% increase in critical vulnerabilities after just five iterative prompts, suggesting that the more you let the model refine its own code, the worse the security posture gets. When these issues compound over time, the result is a higher total cost than traditional development. However, this doesn’t matter when you don’t think long term , but fast fashion instead. Also, none of this is to say that an experienced engineer cannot use these tools well, because they certainly can. The issue is what happens when the same tools are used by someone who does not know what good looks like in the first place, and there is nobody downstream of them who does either. The output passes the basic test of it runs and looks plausible , ships into production, and accumulates the kind of architectural and security debt that surfaces only when something goes very wrong . Note: There are credible voices in the industry, particularly from the AI tooling vendors themselves, who argue that AI-assisted development raises a floor more than it lowers a ceiling. In this view, the median piece of software has always been mediocre, written under deadline pressure by tired humans, copied from Stack Overflow without much thought, and held together by duct tape. If an LLM produces output of roughly comparable quality in a fraction of the time, the argument goes, nothing got worse. We are simply removing a bottleneck. I find this argument partially persuasive, and partially convenient for the people making it. It is true that a lot of software was already not great, but it is equally true that there is a difference between bad code written by a human who at least understood what they were doing , and bad code written by a system that does not understand anything . The first kind can be questioned and corrected, but the second kind tends to compound, because the person shipping it cannot answer why it does what it does. At least for now. Software is not the only place where this is playing out. The book industry is arguably further along, with estimates suggesting that somewhere between ten thousand and forty thousand AI-generated books are uploaded to Amazon ’s Kindle Direct Publishing platform every month, many without any disclosure that a model was involved. In June 2023, the Kindle Top 100 bestseller list was found to contain only 19 books written by humans . Amazon has since introduced limits and disclosure requirements , but enforcement is patchy and authors continue to push back against what looks like a slow flood. Categories that have been hit particularly hard include travel guides (generated guides to cities the author has never visited, with restaurant recommendations that don’t exist), nutrition and health (generated diet advice with citations to studies that don’t exist), and public-domain rewrites (generated adaptations of older books, relying on the recognizability of titles that the actual authors never agreed to). Travel guides in particular have produced a small genre of stories where readers arrive at addresses that turn out to be empty lots, or follow walking directions through neighborhoods that no human would ever recommend. Note: The defense, again, is that the bottom of the book market was always full of filler, that print-on-demand has been around for a long time, and ghost-written business books and assembly-line genre fiction predate generative AI by decades. However, the new thing is the scale at which low-effort content can now be produced, and the speed at which it can drown out the rest of the catalogue. Authors are competing for shelf space against entities that can ship a hundred new titles in a weekend. A 2025 analysis of 65,000 English-language articles published since January 2020 found that a little over half of all new articles on the internet are now AI-generated , and it’s not only the written word that’s being churned out by machines . YouTube has its own version of the problem, where, according to a Guardian analysis, nearly 10% of the world’s fastest-growing channels feature nothing but AI-generated content , and on Shorts specifically more than one in five videos served to a new user is low-quality AI-generated material . Not even the highly creative and (up until recently) human process of making music is immune to this TEMU-fication . Spotify has been removing ghost artist tracks for years, but the practice scaled up dramatically when generative tools made it trivial to produce convincing lo-fi background music in arbitrary volume. The platform has reportedly removed 75 million spammy tracks in a single year , and high-profile acts like the AI-generated band The Velvet Sundown amassed over a million streams before being unmasked. There has been at least one criminal case, involving over $8 million in fraudulent royalties , built entirely on AI-generated music and bot streams. However, that is no reason to applaud Spotify , as the company appears to fight the AI spam only when it’s someone else trying to make money off of it. However, there is a sliver of hope, as engagement with AI-generated articles reportedly dropped by around 40% in 2024, and human-generated content seemingly still gets roughly 5.4× more traffic than AI-generated material in some studies. About 38% of consumers openly express skepticism about AI-created content, and people do still seem to be voting with their attention. Whether that vote is powerful enough to shift incentives at the platform level is a different question, and personally I’m not particularly optimistic, especially given that the platforms profit either way. Let’s take Netflix as an example. From my understanding, the WGA ’s 2023 deal explicitly prevents studios from treating AI-generated material as source material, or from using AI to write or rewrite scripts, and Netflix was seemingly bound by that agreement until at least May 2026. Netflix ’s own Generative AI Production Guidelines also seem to reflect this, stating that AI is permitted in ideation , but that its use should not replace or materially impact work that would otherwise be done by union-represented writers, actors, or crew members, without proper approvals . While that sounds reassuring on the surface, it is, in my view, a delay and not a limit. The same company has publicly committed to going all-in on AI in its production pipeline , has signed deals with VFX automation providers that explicitly put a chunk of the global VFX workforce at risk, and has already used generative AI in at least one of its programs ( El Eternauta ). The trajectory seems to be “use AI everywhere it is contractually allowed right now, expand into the rest the second the contracts permit it, and spin the result as dEmOcRaTiZaTiOn Of CrEaTiViTy” . So here is my specific (and quite possibly wrong) prediction: Within the next five to ten years, Netflix will offer a basic subscription tier whose catalogue consists predominantly of AI-generated or AI-assisted content. We are talking generated procedural shows where each episode is remixed from a small set of templates, generated kids’ content that is vaguely educational and impossible to remember an hour after watching, and generated dramas that recycle plots from existing IP and vibe the rest. For this, the viewer pays the lowest monthly price, while the platform pays nearly nothing in production cost and keeps an enormous margin. The only “upside” for consumers will be the lack of ad breaks, as targeted advertising will quite possibly be injected in real-time into the show you’re watching, seamlessly blending into the storyline without you noticing it, but ultimately still triggering your ape brain to crave a refreshing soda or a sweet treat . Their premium tier, meanwhile, will become the human-made tier. Series with credited human writers, films with credited human directors, and performances by humans whose likeness has not been digitally replicated. The marketing will not call it human-made , because that would be admitting that the cheap tier isn’t , but the price difference will make it obvious. You will pay extra for the same thing Netflix has been selling you all along, except now it is positioned as a luxury. Clearly, I cannot prove that this is what will happen. Netflix ’s own guidelines, as written, prohibit it, and the WGA deal forced a delay. But once the contractual block has lifted, the financial logic is hard to argue with. A streaming service that can produce good enough content for fractional cost will eventually try to. And, mind you, Netflix is just one example. The same logic applies to every other content-distribution business with a subscription model and a margin. If you want to know what the human side of this two-tier world looks like, I think the best existing model is the handicrafts and handmade goods market . By 2025, that market was estimated at roughly USD 987 billion globally, with projections reaching over USD 1 trillion by 2035 . There is data suggesting that U.S. consumers already spend almost a fifth of their money on handmade goods rather than on mass-produced equivalents, and over half of handicraft buyers globally indicate a preference for products that are eco-certified or made from natural materials, going in the exact opposite direction of what TEMU has been doing. What this market shows is that industrialization does not erase the artisans, but pushes them into a different segment. People did not stop buying handmade chairs when factories started making chairs cheaply. While the masses opted for the cheaper, mass-produced items, a small but sustained minority of buyers continued to seek out the human-made version, and over time were willing to pay a premium for it. If the hypothesis holds, software engineering, writing, acting, illustration, composition and the other content-producing professions will undergo something similar. The bulk of the market will migrate to the cheap, mass-produced, generated tier, while a smaller market will continue to value, and to pay for, work that is verifiably the product of a thinking, breathing, opinionated human being. We are already seeing the first signs of this in agencies that explicitly advertise human-only content (at a premium), and in licensing companies flagging tracks as human-composed to distinguish them from AI library music. I think that the interesting question is not whether this segmentation will happen, but what proportion of the market ends up in each tier, and how robust the upper tier turns out to be. There is a darker version of this analogy. Roughly 57-60% of the daily caloric intake of the average adult in the United States and the United Kingdom now comes from ultra-processed foods . Across 22 European countries the share ranges from 14% to 44% , depending mostly on how protected the local food culture has remained. These foods are cheap, abundant, available everywhere, and nutritionally inferior to the alternatives in ways that have been studied at length . People know this, but they eat them anyway, often because the alternatives are slower, more expensive, harder to find, or require skills that have not been taught. I suspect that AI-generated content is on the same path. The cheap tier will not be a marginal phenomenon serving a marginal audience, but it will be the default , the cornerstone of how most people consume software, entertainment, news, and information, because it is what the platforms will serve them and what their monthly subscription covers. Some will care enough to seek out the alternative, but most will not, just as most people, knowing what they know about ultra-processed food, do not change their grocery habits. Probably the strongest counter-argument to all of this is that LLMs are still early, that the quality issues are transient, and that within a few model generations the gap between AI-generated and human-generated work will narrow to the point where the distinction stops mattering or might not even be possible anymore. If that is true, the two-tier picture collapses, because there is no longer a quality difference to justify the upper tier, only a marketing difference. The handmade analogy breaks because, unlike a hand-built chair, a generated novel is functionally identical to a written novel once you can no longer tell them apart. However, I am doubtful that this is going to be the case. There are tasks where I have watched the gap narrow faster than I expected, but there are also tasks where the gap has stayed stubbornly fixed and the failures have just gotten more sophisticated. My instinct is that for narrow, well-bounded technical work, the gap will close further. For long-form work that depends on a coherent worldview, lived experience, and, most importantly, emotions, I doubt it will, because the model has none of those. The second counter-argument is that the consumer backlash will be stronger than I am giving it credit for. The 40% drop in engagement with AI-generated articles is not nothing, and platform incentives may shift if users start to penalize AI-flooded feeds. Apple and others have started experimenting with content provenance and disclosure schemes that, if widely adopted, could stop the worst of the flooding. So it is possible that I am underestimating the immune response . The third counter-argument is, that the cheap tier might not be sustainable at all, because AI-generated content trained on AI-generated content degrades model quality , and the broader ecosystem ends up poisoning its own training data. If that turns out to be the dominant dynamic, the cheap tier could collapse before it becomes entrenched. I think all three of these arguments are valid and have a certain weight to them, but none of them are strong enough, in my view, to make me confident that the TEMU-fication will not happen. They might modulate how it happens, but they probably do not stop it. Initially, I went looking for an optimistic ending for this write-up, to say that software engineering is not going away , and writers are not going away , and actors are not going away . And while all of that is, I think, true, none of it should be confused with things will look the same . What I expect, and what I am to some degree already seeing, is that the people producing software, books, music, scripts, and other human-made work will not disappear , but they will get pushed into a narrower, more specialized, more “luxury” -coded part of the market, pretty much the same way hand-bound notebooks, independent record stores, and small bakeries that mill their own flour did. There will still be a livelihood in it, at times a very good one, but it will look vastly different, and there will probably be fewer people making a living in these fields. My assumption is that they will be more visible inside their niche, but less visible outside it, and they will make their case in part on the basis of provenance , where something was made by a human who knew what they were doing, and you can tell. Meanwhile, the bulk of what most people interact with will, I suspect, be generated. Some of it will be fine, and some of it will be ultra-processed , in the same sense that a frozen lasagna is ultra-processed. It will be functional, calorically adequate food , but it will not be what your Italian grandmother was making. People will nevertheless eat it because it is there, it is cheap, it is convenient, and because the alternatives have been priced out of their daily life. There is no “inevitability” to it, because none of this is really decided yet. There are still choices, made by platforms, by regulators, by consumers, and by the people doing the actual work, that will shape which tier ends up being how big and how durable. The handmade market exists because enough people kept buying handmade goods to make it viable. The human-made tier of software and digital goods will exist because enough people keep buying it, or it won’t exist at all. If you are someone who writes code, or stories, or music, or scripts, by hand, with intent, and with a point of view, I do not think the LLM is going to kill your job . I do think, however, that it is going to change the shape of the market you operate in, push you toward the upper tier (whether you wanted to be there or not) and ask you to make a more deliberate case for why your work is worth the difference in price. For the rest of us, the more interesting question is which tier we are choosing to consume from, and whether we are choosing it on purpose, or just because it was what the algorithm served us by default. I have my suspicions about the answer, but I would love to be wrong.

0 views
Ankur Sethi 4 weeks ago

Prevent cognitive debt by manually retyping LLM-generated code

Despite what I said in April , I'm still using coding assistants on my personal projects. Using them to one-shot entire features leaves me unsatisfied and disoriented, but I do enjoy using them to fast-forward through the boring parts of my projects. However, allowing my coding assistant to roam free in my projects leaves me with a colossal amount of cognitive debt. I might hate the idea of poring over the Django documentation to figure out how to add tagging to my website, but I still fundamentally want to understand how it works. Just because a problem is boring doesn't mean I want to fully offload my understanding of the solution to a machine. Of course, I could review every single line of code the LLM produces. That's what most developers are expected to do in this cursed year of 2026. Robots raise PRs, humans review them. It's a brave new world. But I don't enjoy reviewing AI-generated PRs. Poring over hundreds of lines of overly-defensive, badly-commented, subtly incorrect code is not fun. I might grudgingly do it for an employer—while making sure said employer becomes an ex-employer as soon as possible—but I'm sure as hell not doing it for my personal projects. Personal projects must be fun above all else. The joy of working on personal projects comes from the process, not from the outcome . So what's a boy to do? How do I offload the boring work to LLMs without ceding control of my own work and cognition to the slop machine? I've come up with a solution that's grossly inefficient and perhaps slightly comical: I ask my coding assistant to generate code in the chat, then manually make all the edits myself. I have these instructions in all the agents files in my personal projects: I want to understand every line of code that goes into this project. Never create, edit, move, rename, or delete project files unless I explicitly ask you to do so. Instead, show me every proposed edit in the chat so I can type it in manually. Do not run commands that modify project files, install dependencies, or change repository state unless I explicitly request that action. Instead, show me those commands in the chat so I can run them manually. I'm an experienced developer. Do not explain syntax, APIs, programming concepts, or implementation details unless explicitly asked. Using LLMs this way allows me to work faster than not using LLMs at all, but I'm still slower than those who are willing to allow the machine to think for them. Instead of being 10x faster, I'm probably only 2x faster. But what I lose out on in terms of speed, I gain in terms of a deeper understanding of my code. As I manually type every single line of LLM generated code into my editor, I build up a mental model of how it works and fits into my existing codebase. If I don't understand an API or algorithm, I can stop to look it up, or just ask the LLM to explain it. Typing the code myself forces me to slow down, which means I'm more likely to detect hallucinations or bad design choices the LLM might have made. I can clean up the code as I go, reorganizing it, refactoring it, adding comments, and generally adapting it to my own taste. Most importantly, this workflow allows me to build a spatial map of my codebase. I know where every bit of functionality lives in the codebase. When I need to make a change, I know exactly where I need to make it. It not only helps me work faster within my projects, it also makes it easier for me to better prompt and instruct the LLM in the future. When I was learning to code as a teenager, experienced programmers would often tell me to never copy and paste code into my projects. If I was learning from a book, I was advised to copy all the examples into my computer and make sure I could run them. If I was learning from a blog post or forum answer, I was advised to type it out and adapt it to my codebase so I understood it completely. Manually typing LLM-generated into my codebase feels like the exact same learning process. It might not be the most efficient way to work with an LLM, but I value comprehension over productivity. I've been doing this for a few months now, and it's been working well for me. I plan to continue using this workflow for as long as I can. I fear the software industry is taking on a large amount of cognitive debt that we'll have to pay back very soon. There will come a time when we no longer understand how large parts of our digital infrastructure are put together. I might not personally be able to change the course of the entire industry, but I can at least make sure I completely understand the software I put out into the world. Anything else would be professional malpractice.

0 views
Ivan Sagalaev 1 months ago

Categorization with NLP

Since launching my categorization tool Shoppy I've had some fun analyzing collected data which resulted in considerable complication of the prediction model. And now I feel the urge to write a deeper dive into its inner workings. I'm not sure how useful it would be for anyone who isn't a part of the Grocery Categorization industry, but hopefully some NLP tricks could be at least interesting to any general practitioner. Please note that I'm by no means an NLP expert! Part of the reason for writing this kind of posts is to try and nerd-snipe someone who knows more into sharing their expertise. Oh, and it's not a short one, this! Settle down :-) Just to remind everyone, what I'm trying to solve for my slowly developing shopping list app is suggesting grocery categories for products. So that it knows that "Milk" is dairy, "Apples" is produce and so on. This categorization helps with grouping and sorting, and it also looks nicer with colors and appropriate icons. The usual approach to solving such problems is Machine Learning, and specifically — classification . That doesn't work for me though, as I couldn't come up with any means of collecting enough data myself, and hiring a consultancy is way above the budget of a tiny personal project. So instead I'm manually crafting a clever algorithm with explicitly handled edge cases. The first step is turning free-form input into something more formal and predictable: a set of lexemes. Step by step, it looks like this: This gives me a normalized, stable input key independent of basic morphological variants: A note on stemming. I'm using the original Porter algorithm . It's widely supported, but is also the most simplistic. I don't really care about the correctness of the result from the point of view of the English language. As long as the algorithm used to produce the data is the same as the one used for checking against it, I'm good. The straightforward design for a database is just a mapping from a whole key to a category: . But that would require listing all likely real-world products: all sorts of apples, peppers, beans, etc. Which doesn't work for me since, as I mentioned, I don't have a firehose of data to fill it up. But you'll notice that mostly all such product are defined by one word: an apple is an apple and is produce, regardless of the sort. So let's reify this in the form of a CSV file: These one-word keys are called unigrams . To match a search query, we can simply look up every separate unigram from it in the database. It works for surprisingly many cases, but it breaks on things like "apple juice", because despite having "apple" in it, it's a beverage, not produce. This is fixed by making the order of rows in the database significant, so that goes before . Then, if several of the unigrams in a search query match, the earliest one wins. This might smell to you like something prone to potentially irresolvable ordering cycles, but I actually found that I only really need two groups of significance: Things like "juice", "milk" and "oil" are usually derived from something, as in "apple juice", "oat milk" and "olive oil". Keeping those derivations above raw ingredients ensures those tings are categorized correctly. And don't take the word "raw" too literally. There are things like "ketchup" in there! But since nobody puts "tomato ketchup" on their shopping list, it is considered "raw" in my domain. The next problem is combinations of words. "Spaghetti squash" is not a pasta, but a kind of squash. And "apple sauce" is neither produce, nor a condiment, but a snack! In both of these examples no single unigram is enough to correctly identify the product. This means I need to consider two-word terms called bigrams : All bigrams come before unigrams, as their purpose is to serve as more specific disambiguations of conflicting unigrams. But this ordering is less significant than the groups (Derivations and Raw), so each of those gets their own set of bigrams. At search query time, I produce bigrams as all possible combinations of two from the set of unigrams. Then they're added to unigrams as search terms: The lookup process stays exactly the same: just check all terms one by one. By the way, I need this code to also work in the Kotlin code of my Android app, and since Kotlin doesn't have in the standard library, I wrote a trivial recursive implementation by hand: Felt like solving a coding interview problem :-) A few words about "pepper"… When combined with another word, it most often belongs in produce: "bell pepper", "chili pepper", "serrano pepper", etc. And since there are many of such, I want to classify the unigram "pepper" as produce to cover them all. However the single word "pepper" usually means ground black pepper, which is a spice! I tried to complicate the model at first, supporting wildcards like "* pepper". But I didn't find any other exception that would use it, so I de-complicated the model back and hard-coded a dumb `if` substituting "pepper" with "black pepper" before any lookups :-) Practicality beats purity! Human languages tend to merge words commonly used together. So something straight and forward becomes straight-forward and then just straightforward. I realized I need to care about it here after I collected examples like "redbull" and "lipbalm". Both should properly be spelled as two separate words, but people usually don't consult a dictionary before doing their groceries, so… Funny factoid: "breadcrumbs" and "seaweed" are spelled as single words. This can't be solved by spellchecking because my database doesn't contain original spelling: it has the bigram "bull red" and the unigram "balm", and they both are too far away from those search queries. Instead, I started thinking about splitting words into syllables. To my everlasting surprise, nltk turned out to have several of those, from which I picked SyllableTokenizer . It is pretty simple, but it does work for "redbull" and "lipbalm" as you'd expect. At first, I tried to switch the entire database to contain only syllables instead of full stemmed words, but it didn't work out. There are common syllables like "bar" and "can" that are also full words in their own right. So something like "barbecue sauce" would be split into and will suddenly become a snack because I have a record saying that any kind of "bar" is a snack. So instead I converged on a two-step lookup. First I split the search query into regular words and look them up as I did before. Then, if I didn't get a match, I split the words further into syllables and do a second lookup. The syllabilization algorithm was another thing I had to port to Kotlin, because I couldn't find anything like that in the entire Java ecosystem. Drop me a note if you want that code for some reason. Spell checking proved to be necessary nonetheless. Do you know if it's "fusilli" or "fussili"? Does "camamber(t?) has a "t" at the end? Or is it "haloumi" or "halloumi"? (Answers: 1) the former, 2) it does and 3) both spellings are correct.) Unfortunately, a regular spell checker against English wouldn't work very well: all kinds of international foods and deliberately misspelled brand names kill the idea. Fortunately, I can spell check against my own database, which naturally represents the entire language that my app knows! And the standard approach to spell checking is to employ some "edit distance" algorithm which shows how far apart is the spelling of two arbitrary words. I went with Damerau-Levenshtein (standard Levenshtein plus transpositions). I also had to come up with empirical numbers for maximum allowed distance depending on the word length. This was done completely unscientifically, and I expect to have to tweak it further. But for now it works! The most astute reader might have already spotted a problem: what was a single hash-table lookup before the introduction of edit distance calculation has now turned into a linear search with each test being way more expensive than a straightforward string comparison. So I had to to some low-hanging fruit plucking in terms of optimizations: But I have to say that the main thing working in my favor here is that the data is just small ! I want to highlight two particularly weird wins made possible by the combined effect of syllabilization and spell checking. And I just want to state for the record that the existence of such a thing as "mayochup" made me die a little inside. Is the trouble of mixing mayo and ketchup by yourself in your kitchen so unfathomably hard that you need a brand to produce it for you? Geez… Initially I thought of putting the collected data up on Kaggle , but since it's now heavily dependent on a custom algorithm, I'm just going to leave it as is in the Shoppy repo: Lowercase the input string Normalize it into the [NFD Unicode form][] — the one that separates accents from their characters, which lets me get rid of the former (not everyone bothers typing "crème fraîche" properly). Split the string into "words", ignoring everything else like punctuation and whitespace. In my case words are defined as alpha-numeric characters and an apostrophe (because I want things like "7'up" to be one word). For each word, get rid of apostrophes and stem them. Sort the resulting list of word stems. Assume people don't usually misspell the first letter of a word. Don't bother comparing words that differ in length by more than the maximum allowed distance. Don't bother comparing words with different n-grammness (like bigrams and unigrams). If a search term matches exactly, don't bother two check the rest of them to find a closer match. A single term "may o" works for: "mayo", "mayonnaise", "mayones" and even "mayochup". Syllabilizing these words produces "may" and some variant of "o"/"on"/"oc", which is then smoothed over by spell checking. One of the previously unknown to me items in the feedback was submitted as "separilla", which turned out to be a misspelling of " sarsaparilla ". So I picked out three syllables conveying the meaning: "sar", "pa" and "ril". And, thanks to spell checking, they're enough to cover both the correct spelling and a few incorrect ones. lib.py : all the functions terms.csv : terms database (in a weird format of CSV-with-blank-lines-and-comments) kinds.csv : hierarchy of defined categories (I didn't mention it, but you'll get the idea)

0 views
Justin Duke 1 months ago

Cursed knowledge

Nick pointed me towards Marcin who pointed me towards immich's list of cursed knowledge the other day, and it has already become a running joke in the Slack. Here is a baker's dozen of Buttondown's own cursed knowledge: 1 Yes, that's the joke. The Python library assigns the device family to every non-Mac desktop browser The HTML attribute only filters what the file-picker dialog shows you; drag-and-drop and clipboard paste bypass it entirely. Safari and Chrome re-serialize quoted CSS custom-property strings differently when you read them back via : Chrome keeps the single quotes, WebKit rewrites them to double quotes. Django emits a — which fails our CI — for any cache key over 250 bytes or containing a space or control character. Python's has no default timeout and will, given the opportunity, wait forever. SPF directives recursively chain DNS lookups against a hard cap of ten — exceed it and you get a , which can fail authentication for all of your mail. Outlook and Hotmail enforce mandatory TLS but serve a certificate chain rooting at DigiCert Global Root CA (G1) — a root that Ubuntu has since removed from its trust store. Django's tests whether the key exists , not whether its value is JSON . does not lock rows in the order you listed them — Postgres locks them in executor scan order, which is a wonderful way to deadlock two queries that both thought they were being careful. A postgres cannot exceed ~1MB. Stripe will send subscription update events for paused subscriptions. The Python library assigns the device family to every non-Mac desktop browser The HTML attribute only filters what the file-picker dialog shows you; drag-and-drop and clipboard paste bypass it entirely. Safari and Chrome re-serialize quoted CSS custom-property strings differently when you read them back via : Chrome keeps the single quotes, WebKit rewrites them to double quotes. Django emits a — which fails our CI — for any cache key over 250 bytes or containing a space or control character. Python's has no default timeout and will, given the opportunity, wait forever. SPF directives recursively chain DNS lookups against a hard cap of ten — exceed it and you get a , which can fail authentication for all of your mail. Outlook and Hotmail enforce mandatory TLS but serve a certificate chain rooting at DigiCert Global Root CA (G1) — a root that Ubuntu has since removed from its trust store. Django's tests whether the key exists , not whether its value is JSON . does not lock rows in the order you listed them — Postgres locks them in executor scan order, which is a wonderful way to deadlock two queries that both thought they were being careful. A postgres cannot exceed ~1MB. Stripe will send subscription update events for paused subscriptions.

0 views
Filippo Valsorda 1 months ago

Production ML-DSA Verification in 350 Lines of Python

I don’t do a lot of Python, at least not in my most recent life. 1 However, I happen to have just written a production ML-DSA verifier in pure Python . It’s 350 lines of code (plus many more of tests), it supports all parameter sets, and I am pretty satisfied with it. You can fetch it as from PyPI , thanks to William Woodruff , or you can copy-paste it: it’s a single file without dependencies and it’s dual-licensed CC0 and 0BSD. It works with Python 3.8 and later. The API is modeled after the excellent pyca/cryptography . I hope this will make it easier for some projects to migrate to post-quantum authentication, which has suddenly become more urgent than we all anticipated . In particular, I hope it will unblock some client applications that can’t use C extensions for portability reasons. Modern Python package management , typing , and linting are also a lot more powerful 2 than in the early Python 3 days, and the result is a pretty readable ML-DSA verifier. ML-DSA is actually very simple to implement with its 23-bit base field: we use Python integers (without even needing Python’s big integer support) and SHA-3 from hashlib. There are 86 lines of throat clearing, 27 lines of base field (arithmetic, , ), 28 lines of sampling ( , ), 39 of polynomials ( , ), 25 of NTT, 30 of parsing and packing ( , , ), 35 of key expansion ( , ), and 80 of actual signature verification ( , , ). Performance is… decent? 230 ML-DSA-44 verifications per second without precomputation. That’s 60x slower than Go, but not 1000x. The only optimization change I made was using integers instead of field elements in the NTT hot loop . The implementation is tested with the full reusable ML-DSA testing stack: Wycheproof test vectors and CCTV accumulated vectors , using pytest and muzoo for mutation testing. It has 96% branch coverage, and more importantly it kills every mutation I (and Claude) could think of. (ML-DSA testing techniques deserve their own article.) The project started as a way to double-check the tests of the tests of my Go crypto/mldsa implementation. How do you know your tests are good and comprehensive? You add bugs (“mutations”) and you check that the tests fail. What if you skipped a check though? There won’t be any code to introduce a bug in! The obvious solution is to write a different implementation from scratch, then introduce bugs there, check that the tests catch the bugs, and then port the tests back. Duh. Anyway, pure Python might not be particularly well-suited for cryptography that involves secrets because producing constant-time code could be difficult. However, a signature verifier involves no secrets, and Python is expressive and, most importantly, different from Go, making shared mistakes less likely. You might want to follow me on Bluesky at @filippo.abyssdomain.expert or on Mastodon at @[email protected] , but I can’t promise any more Python. The CENTOPASSI is not all smooth riding, that’s part of the point. However, I am a little annoyed at the local who I had called and who said this road was closed but totally doable on a motorcycle. My work is made possible by Geomys , an organization of professional Go maintainers, which is funded by Ava Labs , Teleport , Datadog , Tailscale , and Sentry . Through our retainer contracts they ensure the sustainability and reliability of our open source maintenance work and get a direct line to my expertise and that of the other Geomys maintainers. (Learn more in the Geomys announcement .) Here are a few words from some of them! Teleport — For the past five years, attacks and compromises have been shifting from traditional malware and security breaches to identifying and compromising valid user accounts and credentials with social engineering, credential theft, or phishing. Teleport Identity is designed to eliminate weak access patterns through access monitoring, minimize attack surface with access requests, and purge unused permissions via mandatory access reviews. Ava Labs — We at Ava Labs , maintainer of AvalancheGo (the most widely used client for interacting with the Avalanche Network ), believe the sustainable maintenance and development of open source cryptographic protocols is critical to the broad adoption of blockchain technology. We are proud to support this necessary and impactful work through our ongoing sponsorship of Filippo and his team. Fun fact, I got started in open source as a maintainer of youtube-dl.  ↩ I feel the same about the TypeScript ecosystem. It’s fun for a week or two every once in a while, but I wouldn’t want to daily drive any of these ecosystems: it’s too easy to spend a whole day updating dev dependencies and fixing linter errors and get the mistaken impression of having gotten anything done.  ↩ Fun fact, I got started in open source as a maintainer of youtube-dl.  ↩ I feel the same about the TypeScript ecosystem. It’s fun for a week or two every once in a while, but I wouldn’t want to daily drive any of these ecosystems: it’s too easy to spend a whole day updating dev dependencies and fixing linter errors and get the mistaken impression of having gotten anything done.  ↩

0 views
Pete Warden 1 months ago

How to set up Raspberry Pi wifi by just talking

As soon as I received my first Raspberry Pi, I knew that it would be a wonderful platform to bring AI into the physical world. Since the initial hardware didn’t have good CPU support for fast arithmetic, I ended up writing code that ran on the GPU so I could get the speed I needed for early deep learning vision models. That was in 2014, and since then the capabilities of both Pis and AI have skyrocketed, and I’m even more convinced that there’s massive potential in combining them. To show you why, I’d like to demonstrate how open-source AI running locally on a Pi has solved some practical problems I’ve run into, and hopefully inspire you to build your own projects using the new possibilities. Pis are great for systems that need to be out in the world, doing specialized jobs. I’ve seen them work well in all sorts of roles, from badge scanners to wildlife cameras. I even run a class that teaches students all about edge AI using the platform. While the boards are generally easy to use, the most frustrating part for the students and instructors is the setup process. While the latest imager makes it straightforward to configure settings like a wifi network to join or enabling SSH when you’re flashing a card, getting the students to the point where they can connect to their Pi using VS Code from their laptop could often take multiple sessions. The biggest problems were: A lot of these issues were solvable if you plugged the devices into a monitor, mouse, and keyboard, but this has its own problems. It meant we needed to provide that equipment to all students during class, and allow them to take it all home too, so they could update the configuration for their personal networks. It also required an extra power socket per student, for the monitors, which added up in a class where we already had to bring in a cart full or power strips. The monitor connections also weren’t always plug and play, we found we often needed to boot with a screen attached to have the display recognized. This isn’t just an educational problem either. One of the reasons that I believe the Internet of Things failed is the setup tax involved in getting smart devices running. According to manufacturers I’ve worked with, less than 30% of their smart appliances ever get connected to the internet, because the process of downloading an app, setting up an account, connecting over Bluetooth, and then typing in the wifi name and password takes too long, and is too errorprone. Even professional installers sometimes struggle with configuration in enterprise and industrial environments. So, what can AI do to help? One of the biggest developments in AI over the last few years has been the development of highly-accurate open-source Automatic Speech Recognition (ASR) models, also known as Speech to Text (STT). OpenAI were the pioneers in this area, releasing the family of Whisper models in 2022. These offered accuracy that was competitive with the models used internally by large tech companies like Google and Apple. These new models allowed startups to begin building voice applications that had never been possible before, and led to a new generation of dictation and meeting note tools like WhisprFlow. One of my dreams as I dealt with all of the configuration issues was a voice-based system that would allow me to simply plug in a headset and set up everything by talking to a Pi. Whisper made this dream seem more realistic, but as I tried to use the models on local hardware, I realized that they were too slow for any kind of interactive application. To address that my startup trained new models from the ground up, designed specifically for realtime applications on affordable hardware. These Moonshine models are smaller than Whisper (our high-end is 250 million parameters versus OpenAI’s 1.5 billion) while offering better accuracy. We also implemented a streaming approach, where a lot of the work is done while the user is still talking, so we can return results even faster. This allows us to return more accurate results than Whisper v3 Large, in just 800 milliseconds on a Pi 5 , whereas even the less-accurate Whisper Small takes over ten seconds. I was excited because this meant I could finally build a responsive voice agent that runs locally on a Pi, something offline-first, and fast and flexible in how it responds. This kind of system needs more than just an STT model, it needs to decide what the user means and respond by taking actions and talking back with a Text to Speech (TTS) system. The Moonshine Voice framework includes modules for conversation flow and TTS, so I was able to use it to build pi-help-bot , a local voice agent for network configuration on the Pi. The application listens to the microphone for commands like “What is my IP address?” or “Help me set up the wifi please”, figures out what actions to take, and responds appropriately by talking to the user. It’s written as a Python script, and here are some snippets that show how it works: This code is a function that uses the netifaces library to figure out the Pi’s address on the local network, so instead of having to connect a keyboard and display or decode the output of nmap, you can ask the question and hear the result, all in just a few seconds. Unlike older voice interfaces, the phrases the user says don’t have to be exactly the same as the one you register an intent with. Instead the framework matches incoming speech against a small, local LLM, so that variations “Hey, can you tell me what my IP is?” work too. This was important to me because one of my biggest frustrations using voice interfaces like Alexa is that they need particular wording to trigger commands, but these wordings aren’t discoverable, so figuring out how to make something happen can require a lot of patience. The IP address command is the simplest kind of conversational flow, where the user asks a question and the system immediately responds. Not all interactions can be handled as simply as this one though. Here’s another example that shows how to implement something that needs multiple questions, answers, and confirmations, connecting to a new wifi network: Hopefully you can follow the logic as it walks the user through providing the information required, but you might be wondering about those yield statements. Those hand back control to the dialog controller while the script is waiting for user responses, so the rest of the application isn’t blocked. The end result is a local voice agent that will listen out for configuration questions and commands, allowing users to set up a Pi for remote access with just a headset. For ease of use, I’ve begun customizing the images I burn to SD cards so that this script automatically starts on boot. This means I can start setting up new devices immediately after powering them on. I hope this gave you some ideas about how a local voice interface could help with problems you face. For further information check out the Moonshine Voice project on GitHub to see full documentation on the library, and please give us a star while you’re there, it helps us keep working on this project. There were different networks in the lab and in the students’ dorm rooms, so it wasn’t enough to hardcode a single SSID and password on the SD card. You need the local IP address of the Pi to SSH into it from a laptop, but it can change dynamically every session. Using “<Pi name>.local” would sometimes work, but some networks didn’t support this kind of lookup, and even if they did it required coordination between the students to avoid name clashes. It was easy to forget to set the configuration so that wifi and SSH were available, and since the instructors didn’t always know what network and password they’d be using in the class ahead of time, we couldn’t pre-flash a bunch of cards to speed up students on-boarding.

0 views
Unsung 1 months ago

“Jokes, art projects, or cruel and unusual punishment”

A fun 19-minute video from commonLuke trying to write a short Fibonacci program in increasingly esoteric languages : = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/jokes-art-projects-or-cruel-and-unusual-punishment/yt1-play.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/jokes-art-projects-or-cruel-and-unusual-punishment/yt1-play.1600w.avif" type="image/avif"> Here are the languages: My favourite was Shakespeare , in which the whole program resembles a play. Believe it or not, but this code outputs “HI”: It was interesting for me to see programming languages that intentionally remove some of the niceties and affordances we learned to take for granted. My guess is most of them are just art, or jokes, or a certain one-upmanship. But I couldn’t help but think of Arika Okrent’s excellent book In The Land Of Invented Languages . The book is about “human” languages like Esperanto and Klingon, but it’s much more interesting than I imagined, and maybe even quite a bit sadder: a story of people afflicted with a certain perfectionism who are not willing to accept languages simply cannot be perfect. #coding #craft #youtube Python, Scratch , and assembly (as control), LOLCODE , a language resembling lolcat memes of yore, Piet , a language whose output looks like Piet Mondrian’s abstract art, Brainfuck , whose programs are made chiefly of punctuation, COW , a version of the above where each of the few instructions is a variant of “moo,” Whitespace , whose code consists only of spaces, tabs, and returns, Chef , where programs resemble cooking recipes.

0 views
James O'Claire 1 months ago

Pip 26.2: –only-deps solves 16 years of app deployment hacks

This has been one of my biggest annoyances working with Python and pip when dealing with projects where that are not meant to be installed as a package , how do you handle dependencies? Think projects like application backends, python scripts, REST APIs etc If you’ve ever struggled with this, you’re going to love this: a PR by Sebastian Höffner opened #13895 will add a new global flag to pip such that you can directly install any dependencies in your without installing the package itself. This is going, for me at least, be a huge boost in the way that I manage and distribute my projects on servers. Vastly simplifying poor manual workarounds that have built up over years. Since it’s been more than a decade in the making, let’s cover the history of poor Python souls stuck trying to figure out how to install dependencies for their scripts or apps. Pip freeze is the classic sure fire first step towards reproducibility documenting exactly which package versions you have installed down to the specific version number. You can then recreate any environment! These dependencies are quite unique to your environment and hardward. Attempting to install from freeze quickly breaks down when you recreate environments on other machines. Different Python versions, OS versions, libraries or machine hardware end up with different requirements of package versions (and full packages as well). This is my oldest memory of working around the issue, it certainly wasn’t the best, but I clearly remembering keeping this around for when I needed it 15 years ago: For me, and likely much earlier others, this sometimes morphed into just manually adding the list of dependencies in a requirements.txt, which I think was a pretty good shortcut. This command has correctly worked the entirety of Pip (2008), and is the fastest way to get pip to install a list of dependencies. These were historically found in Python’s (among others) which predated pip. Dependencies have since migrated to . This works great for libraries and some projects, but it becomes a headache for applications / API frameworks where you may have wrappers running the python code. Editable installs are also not best practice for deployment Installing as a package also created distribution egg files up until 2021, which would become stale if not careful. People, including me, have asked for decades on StackOverflow for how to install dependencies: PIP: Installing only the dependencies (16 years ago) -> Use pip freeze without dependencies of installed packages (15 years ago) -> Use a third party package pip freeze without dependencies of installed packages (10 years ago) -> Use a third party package Is there a smarter way to build requirements.txt files? (2 years ago) -> Use a third party packages or Installing dependencies without the package (1 year ago) -> Use a third party package Look at that train of StackOverflows, Reddit and Python.org discussions. There are hundreds of posts like these over the years, but reading them in order you start to see that the third party libraries were really focusing in on solutions that were more and more useful. After the introduction of , PEP 517 in 2017 added hooks to Pyproject for such as pip, hatch or later uv to use. Another key, which will be used in the ultimate solution, was the 2023 PEP 735 (Dependency Groups) which were introduced to Pyproject.toml to group types of dependencies such as such the user can select which groups of dependencies are needing for a particular install. Finally, we get to the Python ecosystem darling that showed itself to be so useful that it has likely spurred a whole host of changes to Python / Pip that were previously stuck to finally get the attention they deserved. I think it’s worth noting, that while in the posts above there have been may iterations of build tools used for many different use cases, none ever reached the popularity that has achieved. The uv solution: UV crashed onto the scene and took advantage of all the ground work laid previously and showed how much pent up demand there was for build tools with options that were fast and whose user facing CLI solved the real world problems of users. Though pip had similar issues, like #7218 Add pip option to install dependencies , dating back 7 years, issues and discussions always burned out or were eventually closed. After the introduction of and the fast growing popularity of it suddenly started to make a lot more sense. In September of 2022 issue #11440 Add –only-deps (and –only-build-deps) option(s) took hold, and continued to grow, currently with 168 likes. And on April 8 of 2026 Sebastian Höffner opened #13895 Add support for pip install –only-deps . As of now, it’s looking like these changes might make it into Pip 26.2 for the month of July 2026. Höffner’s pull request is adding a global option which will install the dependencies of the project, based on your `pyproject.toml` without installing the project itself. Finally after nearly 2 decades of Pip, the ability to install dependencies for application style projects or scripts has arrived. Looking back at the decade of work leading to this there are so many steps that needed to be taken, by python community peps, by pip maintainers and eve by third party packages. But now that we are here a small but high quality of life change is incoming for Python’s pip 26.2. I hope everyone else enjoys this as much as I know I will.

0 views
Julia Evans 1 months ago

Some more things about Django I've been enjoying

Hello! I’m on a funny journey right now where I’m trying to learn how to make websites in a sort of 2010 style, where I have an SQL database and render some HTML on the backend. It’s kind of an interesting journey because it doesn’t necessarily feel “easy” to me to make websites in this way: I never learned how to do it in the 2000s or 2010s, and there’s a lot I need to learn. So here are some Django features that make building this kind of site feel more achievable than when I was trying and failing to use Go’s standard library or Flask. And I’ll talk about a couple of issues with Django I’ve run into. Previously the toolkit I felt confident with for making websites was: I really liked this frontend-heavy approach for these super simple applications but when I started thinking about making something with a lot of different pages (instead of literally just one page), I didn’t feel so excited about the options I saw that involved a lot of frontend code. So I figured I’d try the backend. Writing a backend-focused site that uses as little JS as possible feels the same to me in a way as writing a single-page JS website that does as little on the backend as possible, even though they might seem like opposites. In both cases I’m just trying to keep as much of the logic as possible in one place. Now for some thoughts about Django! I learned that I can define a “query set” class in Django with a bunch of methods with different statements I might want to use while constructing a query: Here’s how I use it in my view code once I’ve defined what all the methods mean: and here’s how I define the methods: The syntax for defining the filters isn’t my favourite, but I spend most of my time just using the methods, and it feels super readable and nice to use, and it makes me want to look into other query builder libraries in the future. In the past I thought “I know SQL, who needs a query builder?”, but this kind of structure does make it really nice to read. I found an example of someone who wrote their own small query builder in Python that I want to read later to think about whether I would enjoy using a more minimal version of this. There are a bunch of little quality of life filters available in Django templates that are super useful for generating HTML. The ones I’ve used so far are: These are all small things individually but I feel like it makes a big difference somehow to just have them available. I think my favourite template filter is : in this site sometimes we use filters like to decide what’s displayed. that will make a link to the same query string with one change, like this to link to the previous date: Or to remove the parameter: I still really love Django’s automatic database system. It’s amazing to be able to just edit a model to add a new field or whatever, and then Django automatically generates the migration. So far we have done 19 database migrations and I think there will probably be more! It makes a huge difference for me to be able to just easily change the database as my understanding of the problem changes. Django’s documentation sometimes offers the option of using class-based views and inheritance to organize the code in your views. For example I have four views that share a lot of code, and I could use inheritance to manage that by defining some kind of parent class and then having my other views inherit from it. I tried it out and I did not enjoy the experience of using inheritance to share code between views. I switched to using functions instead, sort of how this post advocates, and that was a lot more straightforward. I’ve never had a good experience using inheritance in Python and I don’t think I’ll try to use it again. But I don’t mind using inheritance to use the interfaces Django itself provides: for example if I want to define a query set I need to write something like . I don’t think too hard about it and it seems to work. (as a meta comment: I’ve been working on talking about my programming opinions by just saying “THING does not feel good to me, I prefer OTHER THING instead”. That post I linked to says that function-based views are the “right way”. I’m not very invested in whether it’s “right”, but it’s validating to know that other people feel similarly to me about inheritance) At some point the LLM scrapers discovered our site, and started sending us maybe 10 requests per second. I blocked them which is working for now, but it made me think about what the site’s capacity is. I’m used to writing Go backends where the performance situation is pretty straightforward (usually everything is just fast enough), and a Django site is very different. Some light load testing (with ( ) shows that right now we can serve about 2-3 requests per second (on a ~$10/month VM). It’s tempting for me to go down a rabbit hole where I do a bunch of profiling to figure out what’s slow and try to make it faster (there’s py-spy for that, and py-spy is great and super easy to use, and profiling is fun!) But I really don’t understand what I should expect in terms of performance from a Django site and how I should be thinking about at a higher level. Some things I haven’t figured out yet: I think one thing I’m learning about Django is that because it’s a Framework (tm), it’s easy to accidentally misconfigure it. For example, when I was thinking about why my site was slow just now, I read the django performance docs and I noticed a comment saying: Enabling the cached template loader often improves performance drastically, as it avoids compiling each template every time it needs to be rendered. When I’d done CPU profiling I’d noticed that it was spending a lot of time rendering templates! Maybe this could help me! Clicking through the link, I saw that the cached template loader was supposed to be on by default, but I’d turned it off by accident while trying to do something else. I think this “I turned off the cached template loader by default” things is an example of how I still find the django settings file to be pretty confusing and difficult. I guess I should just be careful when I go in there. After turning on template caching, it seems like the site can now pretty easily handle 12 requests per second or so without using all of the CPU. I have not carefully benchmarked the before and after but it seems like it’s made a pretty big difference. One thing that’s been surprising to me about Django performance is that I’ve always heard the advice “if you have a performance problem, check your database queries! Maybe add an index!”. But I’ve been running into a variety of performance issues (like this template caching thing) that are not because of slow queries, so instead it’s been more useful for me so far to start by running a CPU profile. And since I’m using SQLite, any slow database query problem will show up on the CPU profile anyway. Anyway I don’t want to get too far into site performance. Like I said it’s easy for me to get interested in profiling, but actually I know a lot about profiling and it’s not the most important thing for me to learn about. I might say more about what I’m enjoying (or having a hard time with!) about Django later. Trying to write some shorter blog posts recently. static site generators (like for this blog) static sites that do some fun stuff with Javascript (like this sql playground ) simple Vue.js single page apps with either a Lambda as a backend or a Go backend (like mess with dns ) translating plain text URLs into links, or line breaks into ( ) formatting dates ( ) , which takes a Python dictionary and automatically converts it to JSON and inserts it into the HTML as a tag in a safe way If I have a site that’s going to be getting occasional bursts of traffic, do I want to be able to scale up? Do I want to design the site so that more things can be cached? (and do I really have to? caches are so annoying to get right!) The django performance docs say that Jinja is faster for templating, do I want to think about switching templating systems? Those docs also say “{% block %} is faster than using {% include %}”, I wonder if it’s a big difference and if so why

0 views