Posts in Javascript (20 found)
Martin Fowler 5 days ago

Fragments: August 4

There’s been a fair bit of publicity of the Open AI “rogue agent” that hacked into Hugging Face . This prompted Anthropic to check what their models were up to and, to my complete lack of surprise, discovered three incidents where models had gained unauthorized access to data in other organizations. Simon Wilison concluded : It’s abundantly clear now that running evals of cyberattack potential in models is a spectacularly risky business. Every AI lab needs to pay attention to this. Keeping a close eye on what’s happening in those sandboxes is crucial It strikes me that this is akin to a virus escaping from a laboratory. It makes clear that the model builders are not putting sufficient controls in place to prevent these lab escapes. They are morally responsible for any consequences of this, and that should extend to legal liability too. The bigger concern however is that this same kind of thing can happen with any organization running open-weight models. Lots of labs playing around with dangerous tools and little idea how to contain them. We are sitting in state that Johann Rehberger describes as the Normalization of Deviance in AI . No big disasters have occurred yet, despite all of these worrying signs. But when does our Challenger-moment appear? ❄                ❄                ❄                ❄                ❄ If the sense that we’re in the calm before a storm of rogue AIs worming their way into sensitive software systems isn’t enough, there’s also knowledge that AI is also a financial bubble. Big advances in technology, whether it be railways or the internet, come with bubbles, and those of us old enough to remember the dotcom bubble see all the signs of that now - only bigger. The problem is that bubbles may be obvious, but the way they grow and pop, particularly when they pop, isn’t as clear. The dotcom bubble was widely understood to be one, indeed the chairman of US Federal Reserve talked of irrational exuberance . The trouble is that he said this in 1996, and the bubble took years to grow and burst. Even after the bubble popped, an investor would have experienced an excellent 10% per year gain since 1995. So with that in mind, what to make of the warning signs of this bubble? There are various folks calling out flashing red lights, but I confess I’m not enough into financial and economic analysis to gauge how reasonable these warning signs are, or how seriously to treat the sources pointing to them. Those caveats aside, I’ll mention a couple A substack called “Groundbreaker” calls out a parallel to mortgage crisis of 2008/9 . They say the key indicator of that event was “the second derivative” - that is the point when the rate of increase of prices started going down. The point being that the fuel for this bubble, like many bubbles, was that people believed prices were going to keep increasing, and thus it was good to invest. Once the rate of price increases started slowing, then that was a sign that this confidence was starting to ebb, and an early signal of the crash to come. They see the AI bubble as similar, a credit driven asset cycle, where the assets are data centers rather than houses. The article’s argument seems sensible, but the problem with an argument like this is that it’s all very well to say this flashing red light flashed before the last financial crisis, but it doesn’t talk about how often the light has flashed without a following disaster. Another anonymous Cassandra-wannaby is “Hedgie” a financial X-poster pretending to be an intelligent hedgehog. They noted that Alphabet’s revenue is up, but they are spending even more on capital investments . Much of their gains came from paper increases in the value of their stock in Anthropic, which is highly dependent on the bubble’s continuing expansion. Is this a sign that Google is resting on increasingly shaky financial foundations? Chatting to some of my friends closer to all this, they don’t think Google or Anthropic are the weakest link. They think OpenAI and Oracle are the companies most exposed. We’ll need powerful magnifying glasses to find a suitably sized violin for those companies should they collapse. But is this motivated reasoning? After dodgy sounding anonymous people on the internet, here’s a story from a more trustworthy source giving lots of details on Oracle’s investments in AI , much of it for building data centers that power China and Middle East efforts. “Well-respected A.I. analysts” indicate that Oracle provides over 20% of China’s known A.I. computing power. Doing all of this has created a mountain of debt: Oracle’s debt-to-equity ratio is 500%, compared to 15% for Alphabet. Also on more concrete and less anonymous grounds, there’s been a crash in South Korean memory stocks . Is this a leading sign of a wider collapse? Or should we remember that the late 90s saw five stock market corrections of over 10%, each time recovering, before the bubble finally popped. ❄                ❄                ❄                ❄                ❄ All this talk of rogue AIs and popping bubbles sounds rather dreadful, and John Prideaux made perceptive analysis of this dread risk . Pundits like to point out risks of disaster: A good way to sound smart is to predict that there is a 20 or 30% chance of something awful happening. A p(doom) of 20% is big enough to avoid charges of complacency, but small enough so that you probably won’t be called on it. This is what came to mind when Mr Musk told our editor-in-chief that the probability of ai wiping out humankind was 20%. These are worse odds than Russian roulette with a typical revolver. Anyone who truly believes that should be doing everything they can to prevent the construction of data centres. If they are not, that’s an indication that on some level they do not really believe what they are saying. I grew up with a steady dread of nuclear war, thinking our chances of making it to the end of the 20th Century weren’t terribly good. That fear seems quaint now. Here’s hoping that I’ll feel that way about AI in thirty years time. But meantime, as Eric Evans said in a recent talk: “be nice to your AI, just in case”. ❄                ❄                ❄                ❄                ❄ It’s common to disparage government services, including those on the internet. So I feel compelled to mention an efficient interaction with the government. In this case the credit goes to gov.uk , where I just filled in an online form to renew my electoral registration. The process was quick, and everything was explained clearly. (Gov.uk publishes their Design System , which is worth reading for anyone who is gathering information like this.) ❄                ❄                ❄                ❄                ❄ I had a conversation with a colleague who had used AI to get data out of an otherwise closed package system. The system contained product data for a client, some 6 million SKUs with hundreds of attributes on each SKU. It was our client’s data, but was locked in the package, and the vendor was increasing their prices and made it hard to support new features. The client could copy the database, but the database structure was so complex, they couldn’t make sense of it, and had been working for ten months with limited progress. My colleague’s idea was to use an AI to build JavaScript scripts that scraped the UI. Since the data was presented from the UI, it was in a form that we could understand. It took him a week to extract all the data. I’m hoping we can get a proper description of this story, I think this approach is one that could be used elsewhere. I know lots of people are very frustrated with package vendors locking up their data. ❄                ❄                ❄                ❄                ❄ Any seller faces fraud, and a little industry has sprung up to get fraudulent access to tokens . The idea is to abuse free-trial schemes, play games with chargebacks, and find places that have any kind of open access to inference. The tokens go through a couple of layers and are then sold on to users - commonly done in China. Matt Lenhard’s post includes some tips to limit the abuse, but “the truth is that there’s no clean fix” ❄                ❄                ❄                ❄                ❄ I’ve never had any desire to live in Clacton, but I now find it temporarily appealing. Here’s hoping its residents do the right thing and elect Britain’s first recyclon MP .

0 views
Giles's blog 1 weeks ago

How I use AI on this blog

Inspired by this LessWrong post , I thought I'd write about how I use AI here. This is less in the interest of disclosure, more to provide a snapshot of what I'm doing right now so that I can revisit it in the future and see how it changes. And hey, maybe it'll be of interest to you, dear readers. If I were to summarise my working philosophy in fewer than ten words, it would be: AIs identify problems and I fix them myself. With a very specific kind of exception (which I always flag), the text and code on this blog are human-generated. That's not a moral stand, but more a constraint imposed by what this blog is meant to be -- a place for me to learn in public . Every post here is based on an idea I had, and work that I've done. For many posts -- for example, the large-scale coding projects like this one -- I'll have multiple chat sessions ongoing while I do the work, normally with either ChatGPT or Claude, or sometimes both. The amount of input they have varies, but because the value of these projects is in what I learn when I'm doing them, letting an AI do my thinking for me would make the whole thing pointless -- so I take steps to stop that from happening. AIs are, of course, trained to be helpful, and will often explain things in their replies that I would have better learned on my own through experimentation. I'm generally pretty good at spotting when that happens before I've read more than a sentence or two, though, so I can skip reading that part, scroll straight down to the input field, and ask it to operate in more of a rubber duck mode. With the most complicated projects, where each step has a hard dependency on having got the previous one just right, I do use AIs for code review. Let's say that I've built a model that I intend to extend. I'll test it myself (does the loss go down when training, is it generating plausible-looking results?), but if I want to be really cautious, then I'll run the code past an AI. I'll paste it into a chat session and tell the LLM what it's meant to do, and ask it to check if I've screwed up 1 . Again, though, I make it clear that I don't want it to make fixes -- just to point out any bugs. As things progress with a project, I keep detailed notes. When I'm done, I write them up without AI assistance, getting the post to a level where it might be a little messy in terms of how things are explained, but all of the important information is in there. I read it through and make sure I'm reasonably happy with it, and then it's time for what I've taken to calling the editorial board. I paste the draft post into a fresh chat session with an LLM -- right now, this is normally Claude -- and ask for comments. It already has enough information in its memory of earlier conversations to know that what I'm looking for: places where I'm confidently wrong or other technical errors, places where my explanations are missing a step, or where I'm overexplaining things, conclusions that don't really follow from the results of an experiment, and that kind of issue. It tends to spot a few silly grammatical errors and typos at the same time. An important standing instruction is that I do not want it to rewrite anything. Just as with the code, it should tell me where there is a problem, and let me fix it. We iterate on that for a while until we have something that we're both happy with, and then I feed it to the next LLM -- normally ChatGPT. ChatGPT has a much more pernickety attitude than Claude does. I often use metaphors, and it will generally want me to replace them with mathematically rigorous prose. This is still very useful, though. Sometimes the metaphors aren't flagged as such well enough -- or even worse, there are times when the terms I've hit on for a metaphor happen to clash with a technical term, making what I've written misleading at best. I don't always address all of the issues that ChatGPT raises, as otherwise every post would be a mess of hedges and overexplanations and what-have-you, but I like to get to a stage where I'm comfortable that I have a good understanding of specifically why I'm rejecting the remaining points it raises. One other point where ChatGPT has helped a lot is that it's very diligent about checking supporting materials that I link to. When I was recently about to post an article on running an eval on a model, it followed the link to the training code and spotted a silly bug. It wasn't something that materially changed the eval's outcome, which is probably why I'd not noticed it, but it was something that was important to get right if I wanted later runs of the same evals to be solid. Definitely helpful. With that done, I run it past a cast of other LLMs. The exact set varies over time; for the last few posts it has been (in this order) DeepSeek, Grok, GLM-5.2 and Kimi K3. I did use Gemini in the past, but over time it became less effective and just started complimenting me on the post and suggesting related topics to chat about, which was kind of pointless. I'll wait until the next release and then try it again. Because the Claude and ChatGPT passes have generally got rid of anything particularly nasty, this second group of AIs often don't have much to add. However, occasionally they will spot something the others have missed, or have other suggestions, so it's worth spending the five minutes or so it takes to use them. It also helps to keep me up to date with what the other models out there are like. 2 When all of that's done, I run it past Claude one final time, tidy up any remaining issues, then publish it on a private staging site, and read through it carefully myself. The best time for that final readthrough is after dinner, ideally after a glass of wine; the goal is to smooth out the prose, and remove anything overly formal. To make it as close to being fun to read as I can manage. Once I'm happy, I can promote it to the live site and hit the publish button. That probably all sounds much more complicated than it actually is. A short post will normally go through all of that in half an hour -- less if I skip the full editorial board, which I sometimes do. The longer ones can take an hour or two, but given that they're normally the result of a week of work, on and off, in percentage terms it's not that much, and it's worth it for the polish. So, there's no AI-generated text here, but I do lean on AIs to make it the best version I can of what I have to say. How about the code? Again, the goal of the projects I document here is to learn in public. If I'm learning some concept that is expressed in code, then I need to write that code. So that means that anything non-trivial will be something I've written by hand, with AI input limited to code review -- the same rule as I have for the text. Of course, sometimes there are things I'd like to publish that I wouldn't learn anything by writing. Coding up stuff to chart loss curves, or writing a fancy JavaScript visualiser to show what models' parameters are used for would teach me nothing. So for that kind of thing, I just let the AIs get on with it (and even then, for the parameter visualisation, I hacked a first version together in a spreadsheet to check my understanding, and then tested the visualiser against it). I do always mention in the text when a particular bit of code was AI-generated, though. So, my rule is: if I would learn something by writing the code, I'll write it. If not, I'm happy to delegate to an AI. But even then, I apply one restriction: if it's for the blog, I'll ask for the code in a chat session, rather than using a more agentic system like Claude Code or Codex. This is to add friction. If you're using an agent, it's easy for a task to grow, and what started as a throwaway idea can come to consume more and more time and cognitive space. Keeping it in the chat interface keeps things minimal -- or at least, that's how it works for me. Does that mean that I'm against coding agents? This blog is where I post about experiments I've done and what I've learned. At the moment, what I'm learning is all pretty low-level. How does an LLM work? What factors make it smarter or dumber? It's all pretty hands-on, and involves code that I need to understand. I do other things apart from writing this blog, of course :-) And for that I'm keen on agentic tools; I have an OpenClaw agent to help me run my life generally, and use Codex and Claude Code for projects where I'm trying to achieve a specific goal, rather than trying to learn something. But by their very nature, those are not projects that will wind up here on the blog right now. That might change in the future! When I feel that I have a solid, large enough foundation, perhaps I'll be running experiments that I want to write about, where it would make sense for AIs to handle the details, while I focus on the broader strokes. But that time is not now, so right now, you can be sure that every word 3 , and almost every line of code, was written by hand. Even if I do need the AIs to keep me on track and at least borderline coherent. If you're wondering "why not use Claude Code or Codex", I get into why I avoid them for blog-related work towards the end of this post.  ↩ Current thoughts: ChatGPT, being wonderfully true to form, thinks I should clarify here that I'm not including quotes from models here (like the ones here ) when I say that.  ↩ If you're wondering "why not use Claude Code or Codex", I get into why I avoid them for blog-related work towards the end of this post.  ↩ Current thoughts: DeepSeek used to be close to Claude and ChatGPT, but has been left somewhat behind. I have been hearing rumours on X of an upcoming update, though. Grok used to have a propensity to try to turn my posts into clickbait. It would always want to rewrite things, and would suggest titles that were not a million miles away from "Ten things you never knew about RNNs -- number four will shock you!" Recent releases have been much better, though, and I might experiment with moving it further forward in the sequence. GLM-5.2 recently complimented me on my "science fiction" blogpost that mentioned ChatGPT 5.6 Sol. I probed it a bit on that and it said that because the publication date on the post was in 2026, it understood that I was writing near-future SF. From its perspective, the real date was sometime in late 2024. Surprising! I think the last time I saw that kind of behaviour from a model was, well, sometime in 2024... Kimi K3 is really quite impressive. I will be using it more. I particularly like the way it shows a fairly detailed chain of thought -- something it shares with DeepSeek and GLM-5.2, but there seems to be more depth there. It's a pity that Claude and ChatGPT only show summaries in the chat interface, though I understand their reasoning. ChatGPT, being wonderfully true to form, thinks I should clarify here that I'm not including quotes from models here (like the ones here ) when I say that.  ↩

0 views
Chris Coyier 1 weeks ago

CodePen 2.0

Noting perhaps my largest personal career accomplishment, which is launching CodePen 2.0 . Far more work, believe it or not, than the entire creation of the original CodePen. This isn’t the place to describe every detail of what we did and why we did it. If you’re interested, perhaps our Why 2.0? podcast or the What’s New? page. Instead, a couple of stories from the first week of launch. I was working on a demo with someone I’ve never met before. It started on their (classic) Pen. They needed to import some other JavaScript, so they used 3 Pens and pulled in the JavaScript from the other two into the main demo. They also needed an npm package. I forked the Pen and invited them as a co-editor, so we could both work on it together anytime. I moved the JavaScript into files on the main Pen, as that’s much easier to work with. The npm package is in the file for easy version management. We both cleaned it up to our liking. The Keyframers (David and Shaw) got back together and did a live stream on launch day. They also used the invite feature and live collaboration . They worked together for hours, and while there was a bug or two, it was nothing super major, and it went great. One of my favorite bits was that they shared the Live View of the Pen, so as they were working on it, we could play with the demo ourselves. As I was working on the emails we were going to send out about the launch, I was building them in the special language for crafting them: MJML . I went ahead and added MJML as a block to CodePen so I could just build them right in CodePen. Works great , even for weird stuff . Many more Blocks to come. I friggin love how I can make little websites and deploy them right through the Pen Editor. Like the one for our slideVars library or codepen.school . It just makes me wanna build a ton of little weird websites.

0 views

Read This Before You Buy That TV Streaming Stick

Security experts have been sounding the alarm for years about the risks of using generic TV boxes that promise unlimited content streaming for a one-time fee, warning that they secretly rent the user’s Internet connection out to strangers. But a groundbreaking new analysis finds these devices also routinely spoof themselves as mobile phones clicking ads on AI-generated websites as part of a sprawling operation that seeks to defraud online merchants and advertising networks. Pedro Falé is a threat researcher with the security firm Bitsight . Falé told KrebsOnSecurity he was able to peer inside a vast and complex ad fraud network by registering an expired domain name that was used to coordinate fake ad clicks across a particularly popular brand of these streaming devices known as H96 . An H96 TV streaming device currently advertised for sale on Amazon. Falé said the domain he scooped up was previously used for telemetry, periodically collecting full hardware information and the entire list of installed apps from tens of thousands of H96 streaming sticks plugged into television sets around the globe. But upon inspecting the traffic being funneled to the domain, he discovered nearly all of the TV boxes transmitting data claimed to be mobile phone models from a variety of manufacturers, including Samsung, Vivo, Huawei, and Xiaomi. “We noticed something was wildly wrong,” Falé said. “Multiple devices reporting to this factory Android TV Box backdoor were ‘phones.'” Image: Bitsight. The researcher found all of the devices reported having the same two apps installed, and that those apps were made by a company called Zhejiang Fengwo IoT Technology Ltd , an entity founded in 2019 in mainland China which operates an ad-publishing portfolio under the name Fengwo Group . Further investigation into the Fengwo Group revealed it has registered multiple patents that match the inner workings of these apps. “Bitsight TRACE identified several Hong Kong, Singapore, and single person ‘legal’ shell identities used to collect the monetization and traced the operation back to a mainland China company known as Zhejiang Fengwo IoT Technology Co., Ltd, which operates under the Fengwo Group,” Falé wrote in a report released today about their findings. Falé said an analysis of the apps shows they help to coordinate an ad fraud network that uses these H96 devices as a captive traffic source to click on ads at AI-generated websites operated by the Fengwo Group. Bitsight discovered the websites contain machine-generated news articles and graphics across a range of categories, including finance, health, education, gaming, music and food blogs. But they also found none of those sites displayed ads unless the device visiting the page matched the spoofed mobile profile of these H96 devices. The domain for the Fengwo Group — fwgcloud[.]com — claims the company is “redefining the boundaries of human-AI interaction,” and that it has created more than 120,000 “AI digital humans” available to rent for everything from emotional companionship to 24/7 customer service and creative design. The homepage for fwgcloud dot com. Falé said the Fengwo Group’s domain shared its SSL certificate data with other domains associated with the apps found on H96 devices, specifically the phone spoofing mechanism. He noted the domain also has an internal wiki platform that directly ties the Fengwo Group to a proprietary implementation of a Google-built visual programming language called Blockly , which was originally designed to help kids learn how to write software. According to Bitsight, the Fengwo Group’s employees use Blockly to build the sham websites, allowing low-skilled operators to drag blocks of code together in their Blockly editor — without any need to understand what the underlying code blocks do or how they work. The Blockly homepage. “An operator can drag blocks together in their Blockly editor, to define each fraud routine, given a task type,” reads Bitsight’s report. “Once the routine is saved, it gets exported as JavaScript and uploaded to the S3 buckets. An operator doesn’t need as much understanding of the underlying technicalities, as it is all set in place for ease of use.” Bitsight even found one of the Fengwo Group app developers mentioning exactly these advantages, noting the developer remarked that “only a small number of highly-skilled developers are needed to build the template execution-unit images,” and that “developers who create execution units from those templates have significantly lower technical requirements, greatly reducing the company’s operating costs.” Falé said if a user’s H96 streaming stick is selected for a specific fraud task, it will be pushed the appropriate Blockly module according to the task desired, which can include silently launching a web browser, visiting websites, browsing pages, managing tabs, and clicking on ads. To ensure the TV boxes masquerading as mobile phones can reliably click on ads displayed via the AI-generated websites, the Fengwo group “fuses three vision and reasoning systems into a single interface,” allowing the bots to correctly identify an ad on the webpage and navigate the site much like a human would, the Bitsight report observed. Examples of ad landing pages linked to the Fengwo Group. Image: Bitsight. Bitsight found the H96 devices were either relaying residential proxy traffic or participating in ad fraud, but never both at the same time. In fact, they concluded that when these TV boxes detect an HDMI signal from an attached television — indicating the user intends to stream video content — the box is usually functioning as a residential proxy. When the TV is off, it switches back to waiting for ad fraud jobs. Falé said he believes the TV boxes are set up this way because its ad fraud activities are far more resource intensive and could interfere with the device’s stated purpose — streaming video content over the Internet. Despite repeated warnings from the FBI and security industry leaders about the security and privacy risks of using these streaming devices, major e-commerce providers like Amazon, Best Buy, Newegg and others continue to sell hundreds of different models and brands that bundle unofficial versions of Google’s Android operating system and are frequently marketed ( via online influencers ) as a way to access a broad array of streaming services and live broadcasts without a subscription. Image: fbi.gov. In addition to enlisting the user’s TV box in ad fraud networks, these off-brand streaming devices almost universally come with residential proxy software pre-installed. This software rents the user’s Internet address out to anonymous paying customers, who run the gamut from aggressive content scraping firms to ticket scalpers and outright cybercriminals. What’s more, because these generic (and generally dirt cheap) TV boxes are all horribly insecure by default and bereft of any kind of authentication, installing one on your home or office network only invites further mischief. In January, the proxy tracking service Synthient documented how multiple botnets had rapidly enslaved millions of TV boxes using a complex interplay of security vulnerabilities in both the residential proxy software and the streaming devices themselves. Bitsight said it tracked approximately 38,000 TV boxes globally phoning home to the expired Fengwo Group domain, and based on that number the report estimates this ad fraud network brings in revenues of close to $50,000 a day (not counting substantial revenue from the residential proxy side of the business). However, Falé emphasized that these estimates are highly conservative and based on telemetry from just one of the Fengwo Group’s core (but older) domains. As for the Fengwo Group’s claim to have 120,000 “digital humans” at their disposal, Bitsight’s report concludes it could be just a clever marketing scheme and/or a way to avoid drawing suspicion to the company’s operations. “Historically, when dealing with proxy services or DDoS, we sometimes see these websites undertake inconspicuous facades, so as not to advertise their DDoS capability or botnet size,” Falé wrote in the report. “This could also be the case here.” If the Fengwo Group truly does have tens of thousands of “AI humans” at its beck and call, it does not appear to have dedicated any of them to fielding inquiries from its own website. KrebsOnSecurity sought comment from the Fengwo Group by emailing the contact address listed on the company’s homepage, but the request bounced back with the reply, “Your message couldn’t be delivered to postmaster@fwgcloud[.]com. Their inbox is full, or it’s getting too much mail right now.” As Bitsight’s analysis shows, when it comes to TV boxes and streaming sticks, it’s best to stick to name brands from reputable manufacturers, and then to be sparing and careful with any apps you choose to install on the device — as many of those can bundle residential proxy software as well . Google says consumers can confirm whether or not a device is built with the official Android TV OS and Play Protect certification by following these instructions . Additionally, Synthient maintains a running list of IoT devices that have been known to ship to consumers with residential proxy software and other malicious apps pre-installed. Careful readers will notice Synthient’s list includes other IoT devices apart from streaming sticks and boxes: As the FBI has warned, residential proxy software has also been found in other popular consumer IoT devices from random brands, particularly digital photo frames.

0 views
Den Odell 1 weeks ago

Your SPA Is Leaking Memory. Soak Test It

Memory leaks are a constant battle for backend teams. A server stays up for weeks, responding to requests the whole time, and if any part of the code running on it has a memory leak, even a small one, that server will eventually run out of memory and crash or restart. So how do these teams know their services won’t end up like this? They soak test them. In a soak test, a team points a script at their server and has it send fake traffic for hours at a time, sometimes thousands of requests a minute. These tests are often automated to run overnight, while the developers are away, and they compare the service’s memory at the end against the baseline from the start. If the memory climbed while the test ran, there’s a leak somewhere in the code, so the test fails and the team has to find it and fix it before the service goes live. Frontend code never used to have this problem, because clicking a link to a new page destroys the memory used by the old page. On a page that only lasted minutes, any potential memory leak was gone before it could grow into a problem. But the web has changed a lot in the last decade. Single-page web apps (SPAs) give you an experience that feels more like a native app than a website. It’s smoother to use, but it means the page is never reloaded and nothing forces a full reset of its memory any more, so if any part of the frontend code running on it has a memory leak, even a small one, that browser tab will eventually run out of memory and crash or reload. I know of teams who force a hard reload of their SPAs every few hours just to avoid this. Electron apps work the same way, along with anything else built around a long-lived web view, since the page underneath is never reloaded either. A static analysis of 500 popular React, Vue and Angular repositories , published in early 2026, found that 86% of them set up a listener, timer or subscription somewhere and never remove it. So how do you know your SPA won’t end up like this? You soak test the frontend too. Gmail was doing this over a decade ago , running memory checks in its pre-release tests for hours at a time, after leaks left some users reporting processes over 10GB. Unless you’re deliberately running something like Meta’s MemLab , your existing tests probably aren’t set up to do this for you. Your Playwright end-to-end suite is the closest thing you have, since it actually clicks around the app, but it still starts each of its tests with a new browser context. It starts from the same place every time and finishes quickly. That’s what you want from a test suite the rest of the time, but a leak needs longer than one test to get big enough to measure. Someone with the app open all day works through the same screens over and over. Detached nodes stay in memory, kept alive by listeners still attached to them, while timers keep firing and the cache keeps growing. To make a frontend soak test, you construct a user flow yourself and run it on a loop, all inside a single browser context. We’ll use Playwright for this, since it’s probably already running your end-to-end tests. The flow starts and finishes on the same screen, so each pass leaves the app where it began. Your simulated clicks act like fake traffic, and since Playwright clicks as fast as the app can keep up, a few hundred loops take minutes rather than hours. If the app is back on the screen it started on, its memory should be close to where it started too. You’re watching for memory that keeps climbing loop after loop. The memory only comes back to where it started if the flow is a round trip, like opening a drawer and closing it, or filtering a table and clearing the filter. Some apps are meant to use more memory as they go, of course, so not every flow can be a soak test. Scrolling a feed that loads more as you go is supposed to end heavier than it started. A chat interface might be too, if you don’t delete the messages after they arrive. Chrome will tell you how much memory the page is using over the Chrome DevTools Protocol (CDP), which is what DevTools itself uses. That limits this to Chromium browsers, since Playwright can only open a CDP session there. We collect garbage twice, for reasons I’ll come back to, then ask for the page’s metrics, which come back as a long list including the heap size, the DOM node count and the listener count: The heap size moves around between runs whatever your app does, so it’s the node and listener counts we’ll assert on. Now we’ll add a function around it. It takes the flow you want repeated and runs it 200 times against a single browser context, taking a reading after a short warmup and another at the end: The first time the drawer opens, the browser has to fetch its JavaScript and the app has to fetch its data, and both stay in memory afterwards. That only happens once, on the first loop. If you take the baseline before it, the heap jumps between your two readings and the test fails, when all that grew was the code and data the drawer needed. So runs five loops before it takes the baseline. The node and listener counts don’t need the warmup. They climb by the same amount from the first loop even on an app with lazy routes and a query cache, because lazy-loaded code and cached data live in the JavaScript heap, while the counts go up when the page gains a DOM node or a listener. Five loops take a couple of seconds, so I leave them in anyway. The second call is there to make the node count reliable. On a React app I tried this on, one pass left a detached drawer in the count on about one reading in six, so a healthy test run would fail. I found a plain page with pure JavaScript and no framework always came back clean after one pass of garbage collection. That leaves us with a test that’s just the flow and the assertion: Leaks normally get found by someone taking heap snapshots and reading through them, which is painful and slow. A node count is just a number, so a test can compare the two readings for you. Readings vary between runs, so this belongs in a nightly job rather than on every pull request. Where the leak involves a listener, that count is the one to assert on, since it only goes up when your code adds a listener, and down when it removes one. The node count catches leaks that hold on to DOM with no listener attached, and a fixed allowance works better there than a percentage, because a drawer sitting at 33 nodes between passes makes one stray node look enormous, while the same node against a 2,000-node baseline is nothing. I used in my code because it’s a nice, round number, above the jitter I saw in initial results and far below what even a small leak would add across 200 loops. That soak test still misses the biggest category in that 500 repository scan, though. Timers left without being cleaned up made up nearly 44% of everything it found, most of that . They sit on the browser’s clock, so a polling check set to run every 30 seconds fires twice a minute, and after 200 loops in two minutes it has fired four times, where an hour of real use would have fired it 120 times. You can work around this mismatch by faking the browser clock. Playwright can replace everything the page uses to tell the time, from and to timers and animation frame callbacks. Installing it before the SPA loads lets the page start up normally. on its own leaves time flowing, though, so you pause the clock once the app is up, and from there it only moves when says so. Advancing 18 seconds on each of the 200 passes adds up to an hour of timers across the run: The clock only fakes timers, so a still takes as long as it takes, and whatever comes back is stored by your app and stays in memory. A poller often calls again once each response arrives, so requests never overlap when the server is slow: This next bit is fiddly and it took me a few goes to get straight. There are two clocks running, a fake one that only moves when tells it to, and the real one, which keeps going the whole time. fires the pending timeout, runs until it hits the , and the call returns while the request is still out. The response lands during the next pass, in real time, and then schedules its next timeout from wherever the fake clock stopped. Your API speed sets the polling rate now, not the 30 seconds you asked for, and across the run that works out at roughly 100 requests where an hour of real use would make 120. The fix is to mock the network as well. We answer each request ourselves, then advance 30 seconds at a time, waiting for the response before the next tick: That covers 200 rounds of clicks and 100 minutes of polling, with every response landing before the clock moves again. What you send back wants to be close to what your API actually returns, ideally identical. If you return 200 bytes where the real endpoint returns 50 kB, the leak in your test is hundreds of times smaller than the one in production, and the test passes. For web sockets, does the same job, letting you send messages to the app as fast or as slow as you want, set up before you navigate just like the clock. Streamed responses, like those used in AI interfaces, are the one awkward case, since only takes a string or a buffer, so you can’t send the response in pieces at a speed you choose. Those counts tell you there’s a leak, but not where it comes from. For that you take a heap snapshot in the Memory panel of Chrome DevTools, then type into the class filter, which leaves you with the DOM your app removed from the page and still has a reference to. Clicking one shows its retainers underneath, so you can see which listener or variable still references it. A soak test is one flow, repeated a few hundred times, with the DOM node and listener counts read before and after. Faking the clock and the network lets it run in compressed time, so a short test run covers hours of real use. You get an answer straight away, and you can leave it running overnight with the rest of your automated tests. This is what I mean when I talk about Fast by Default . You write the test before anything goes wrong, so the problem shows up, and gets fixed, before it reaches production. Most teams find out about a memory leak the day a customer says the app goes slow after leaving it open a few hours. Now you have to find it across the whole app and its Git history, when a soak test would have failed the night someone committed the leak. Going by that repository scan we saw earlier, most codebases leave a listener or timer registered somewhere and their development teams have no idea. That pattern, where performance problems only show up once users hit them and someone drops everything to patch it, is what my book Fast by Default: Practical Performance Engineering is all about fixing. This soak test is one small example of it. The book applies the same approach to loading, rendering and everything else users wait on, and argues for performance being something the whole team owns rather than one person’s job. It’s in early access now, so the chapters are going up as I write them.

0 views
Hugo 1 weeks ago

My Software Factory in the Age of AI

I'm creating this page to document my "software factory". It will be more of a reference page than an article and I will reference it on the resources page of the site. Context : several applications in monorepos ( hakanai , Writizzy , Bloggrify ), polyglot (Nuxt, Kotlin, JS), solo dev, code written largely by agents, continuous deployment to production. ::toc{open="true"} :: Even though the latest generation of agents are now capable of producing quality code (much better than the majority of human developers), code production represents only part of what I call software quality. The rest includes: Part of these risks is resolved with a good understanding of your market, benchmarks, interviews, manual testing, mockups. All of this is also part of my "software factory" even if I don't describe everything here. The code produced is now almost 100% generated, but that doesn't mean it's vibe coding. Vibe coding as defined by Karpathy was experimentation and letting yourself be carried along by a dev session. Here, I'm going to talk about context engineering. The goal is to provide all the necessary context, at the right time, so that the software matches an intention and is systematically controlled. Even though I don't write the code, I'm responsible for it and I need to maintain control over it. The tooling described here answers several questions: Here are the files read by agents before starting. Be careful, the size of the context must remain controlled. Too large a context costs money and degrades the quality of responses if it becomes too heavy. We want to save tokens and optimize when information is loaded. The separation matters: permanent context stays short, specialized context loads when it's useful. Long-term constraint as context. Concrete example: I plan to open-source part of the code. The rule "code intended for open source should never depend on proprietary code" is written in the rules and verified by a test (see layer 4). Writing a future constraint in the context avoids paying for a refactoring later. You can also find constraints on: A constraint is specific to a project and a person. It's not a matter of software quality in the strict sense. It's not better or worse to do feature flagging for example, but it's my preference for partial deployment and feature activation in testing. There's no point in telling an AI to "write quality code", it doesn't make sense. You need to make your own constraints explicit. Example of a rule that allows deferred loading depending on context: It's simple, but it avoids loading the entire skill in a rule (which would be loaded systematically). Here's a more complex rule: Here there are two important things: A skill is a procedure written once, replayed identically. The agent loads it itself when the context matches. This automatic loading can sometimes fail. In that case, you need to explicitly ask to use the skill. The criterion: am I repeating myself? If I explain the same thing a third time, it becomes a skill. I have about thirty, grouped by family: Two things I've learned: Sub-agents : for tasks that generate a lot of reading without much decision (audit, broad exploration, doc writing), I delegate to a sub-agent. It consumes its own context and gives me a conclusion, not a dump of files. I use them less and less, recent agents make their own fairly targeted delegations. You often hear that an AI is non-deterministic and can make mistakes on trivial things that need to be deterministic, like calculating 2+2. That's largely false now, and an AI is no longer "just an LLM": it has many tools to control output. Nevertheless, the best way to ensure a form of reproducibility is to delegate to tools whose job it is. Compilation, test execution, linters, all of that is delegation. You can delegate to MCPs, or to skills that use themselves a command-line tool (CLI). I try to avoid MCPs which consume more context, but I have a few anyway. The repo is indexed in a graph (symbols, relationships, execution flow). This allows you to measure impact levels and find all links with the code being modified: The real issue isn't speed, it's detecting all side effects of a modification. I use two levels of memory: What we want with these tools is to avoid repeating mistakes, document decisions, and not start over with an empty session each time. Claude Code's internal memory mechanism has improved and becomes more relevant than before with the latest versions. However, you need to control it and not hesitate to ask it to delete rules it creates on its own which are sometimes a bit silly. Claude-mem, I honestly have a hard time measuring the negative or positive impact. I don't have enough perspective on it yet. A wrapper (here RTK) prefixes shell commands and only returns what's necessary. It remains very limited, only available for a few tools. The gain is sometimes canceled out because Claude runs the command twice. It doesn't hurt, but I think there's still room for improvement. Note that Claude builds its own tools on the fly in Python or Bash, and also knows how to use filtering mechanisms with , , etc. to optimize the outputs of the tools it uses itself and save tokens. I use two skills for incident resolution: I could also mention the Stripe MCP, read-only as well, which allows in certain specific cases to investigate Stripe configuration issues. Guardrails prevent the randomness associated with understanding and executing an instruction. It happens that an LLM ignores a rule. You need to provide tools that execute automatically. Scripts triggered by the agent harness, not by the agent itself. Other uses that work well: block editing of generated files, require a test alongside any new module, forbid a dangerous pattern. Structuring rules become tests that break the CI. Real example, I have an architecture test that preserves the boundary between future open source code and the rest: It's part of the automated tests, so it can't be bypassed, unlike a rule. I use several things: Each app has its workflow on GitHub Actions. The deployment job has a on the quality job. Nothing goes to production without passing the gate. It's essential in general, even more so for automatically generated code. I shouldn't need to re-explain this, but just in case, I have several types of tests. I give Claude instructions to explain the test hierarchy, gitnexus tells him what to replay to validate these modifications, and he has instructions to write them when he adds/modifies code. As I said in the intro, the goal isn't just to produce code, it's to produce code that serves a purpose. The process often starts with a spec, then a design, then an implementation. Spec first. I have a folder of numbered specs, one per functional domain, indexed in the permanent context. Two skills frame the cycle: one for writing the spec and its plan, one for closing it by updating it with what was actually built. This spec and the discussions can rely on a 'product-marketing-context.md' file that I create in each project and which summarizes my personas, my competitors, my positioning etc... The design. By design I mean two things: technical design and interface. Technical design is part of the spec phase. Most of the time a spec is enough, but some tricky cases require a spec dedicated to a technical component or a technology choice. On the other hand, the design/mockup phase is separate. I do it in Claude Design. I make a functional mockup and iterate until it's perfect. I verify understanding, labels, usability. Then I can pass the result to Claude Code. The implementation. Claude starts from the spec and mockup. He follows the plan made in the spec phase. The spec is meant to be delivered in stages (protected by feature flag). This allows me to do several small implementation sessions rather than one large session, which tends to degrade in quality if it gets too full. The closing step is important: without it, specs become obsolete in six months. Explicit rule in the context: if a spec is vague or inconsistent with what exists, the agent should ask the question, not guess. Progressive delivery. I work trunk based. However I use feature flipping and gating, which represent two different things: I have skills that explain the difference, so agents don't do anything wrong and respect my working process. If you're starting from 0, the first step is the quality gate if you don't have one. You need a control mechanism that runs tests, linters etc... Then start light with a Claude.md that describes the essentials, the why. Add rules as you go for important architecture patterns. As soon as you see procedures that come back often, document them as skills. And then equip yourself with cli and MCPs to interact with your main tools, jira, sentry, etc... Be careful, any skill, MCP, code taken from outside must be scrutinized. These are dependencies that can be vectors of attack. You need to take into account that the technology is still very young, February 2025 if we consider agentic programming. Tooling is improving but you also need to constantly review it. Mid-2025, some instructions in a Claude.md made sense, for example "write a test for each new service". Today it's noise and Claude does it naturally. So you need to be wary of your old rules, sometimes they're obsolete and create noise. I have no way to measure and know if an old rule has become obsolete. Latest versions of Opus are increasingly autonomous. AI takes the initiative on its own to build, look at produced content, read dependency code to understand calls, find bugs, run tests in the browser. It's almost creepy and more rigorous than 99% of humans. Let's be honest, I'm increasingly less useful in implementation phases, but I don't want to lose control of the produced code. I'm torn between satisfaction at having an increasingly efficient software factory and the risk of losing knowledge. I need to find a way to control designs a posteriori, to appropriate the result. I recently added a boyscout.md rule But I find myself having endless sessions. I think I've never worked on a codebase that maintains itself as much. I've rarely improved the product at this level of detail. But it comes with a cost, cognitive overload. I think I'd rather automatically note these elements and categorize them in an online TODO list (trello, todoist etc...). I think the workflow for maintenance should move elsewhere, and be partially automated. I still copy and paste my skills/rules etc… from one project to another. And sometimes it's dependent on my station based on a skill installed locally. I need to find a way to package my skills to deploy them where relevant and to centralize maintenance. In other pending improvement points: the intention (why, for whom) does it satisfy the 4 risks identified by Marty Cagan: Value Feasability we can add: performance and reliability What does the agent know? (context, memory, code graph) What can it do deterministically, without improvising? (skills, procedures) What stops it when it makes a mistake? (hooks, architecture tests, quality gates) marketing copy (labels in the application) how to write migrations types to favor for database fields asynchronism patterns etc. (list far from exhaustive and enriched regularly) the activation pattern, which specifies that this rule only loads if you touch the directory the table that lists all available skills, to be opened only if needed. If the AI doesn't make schema changes, there's no point opening the skill "Multi-file procedure" skills are the most profitable. Example: adding a block to the content editor touches three different rendering surfaces. Without a skill, the agent systematically forgets one. "Up-to-date doc" skills are very important on code that evolves quickly. It allows an agent to understand the entry points and intentions of a feature. before modifying: blast radius, callers, risk level before committing: did I only touch what I wanted to? find an execution flow rather than grep a function name rename via the call graph rather than find-and-replace Claude's internal memory Claude-mem , which allows capturing decisions between sessions a Sentry skill to read information on Sentry and retrieve stack traces a skill that gives read-only access to the database (read/write access is possible directly via on the Docker container in local dev) no file on the "open" side references the "proprietary" side every production file belongs to one of the two sides ESLint for syntax ast-grep for architecture decisions , for example forbidding any call to without going through the OpenAPI client typecheck for typing What matters must be executable. An instruction is followed "most of the time", but it can be forgotten. A hook or test is followed all the time. An error must be documented. Every error must be recorded in a skill or in memory. Context is a budget. Wrappers, filters, sub-agents, summaries: everything that reduces noise keeps reasoning on the problem. Measure impacts before and after editing. We want to avoid the effect "1 bug fixed, 10 produced". Repetitive procedures don't improvise. One skill per procedure. Spec documentation dies if its closure isn't in the process. You need to plan the update and maintenance step. my dependency on Claude. I want to test open weight models but I don't have the hardware for it. Moderate risk in my opinion, the entire ecosystem is moving upward the IDE is becoming obsolete compared to this new workflow. I still use Intellij but I no longer find it suited to our time. I haven't seen an interesting alternative yet.

0 views

The Difference Between a Button and a Link

Of the three proposals in the Triptych Project , my multi-year odyssey to add a few small-but-powerful features to HTML, the one that generates the most questions is Button Actions . The proposal itself is very straightforward: we want to add the and attributes to the button. Button Actions are such a simple primitive that people often ask why they’re needed. The answer rests on a distinction that web users intuitively understand but rarely have to think about directly: the difference between a button and a link. I added a detailed “Buttons vs Links” section to the proposal, but I think it deserves a blog-style explanation as well, because most of the existing ones miss the mark. Links represent a destination while buttons represent an action . Functionally, this means that links let users control what context they open in, while buttons don’t. Web browsers offer countless affordances for re-contextualizing a link. Clicking or tapping the link will navigate the current page to that destination. Mouse users can middle-click the link to open it in a new tab or hover over the link to see where it goes. Context menus (right-click on desktop, long tap on mobile) have lots of link-specific options. Web users are very familiar with the features that come with links. They know how to open them, copy them, bookmark them, share them with friends, and maintain them in an inadvisable number of browser tabs. The semantics of a link—the notion that they represent an independently-navigable destination—make it possible for browsers to build all these features. The hyperlink predates the invention of the browser tab, but when browsers added tabs, websites didn’t have to do anything to support them; links represented destinations that could be re-contextualized, so browsers could simply invent a new context for them to open in. Every website instantly got upgraded with a huge new feature. Buttons have none of these features. By default, they cannot be middle-clicked, control-clicked, or hovered over for more information. Buttons don’t allow you to copy their the way you can copy the of a link. Their context menus contain no affordances for saving the action or doing it somewhere else. These are not omissions, but deliberate choices based on the button’s semantics: buttons trigger actions inside a specific browsing context ( almost always the current one ). Copying, sharing, bookmarking—these are all features for re-contextualizing the action of a link. Buttons serve a complimentary purpose because they don’t allow for any of that. A common misconception is that links are for navigating the page, while buttons are for everything else. This is incorrect on both counts. Buttons regularly perform navigations. Clicking a logout button navigates the current page to a logged-out one; clicking a “search” button navigates the current page to the query results. Both of these are navigations in the HTML standard . They change the URL, they get logged in the session history, and they load a new page. And links are often used in situations where they don’t trigger navigations. Relative links can jump around the current page; mailto links can open email clients; download links can save a file to your computer. None of these are navigations, but they are all “destinations” that can be opened, saved, and shared in customizable ways. Navigations should be represented as buttons when their action happens in a fixed context that is not available to be re-contextualized (e.g. bookmarked, shared, middle-clicked, etc.). A frequent place this comes up is with forms that let you edit something you’ve already saved, like a comment on a website. When you click “Edit”, the website shows you an editable text area with options like this: Users will easily intuit what each button does: Should “Cancel” be a link? No! Its job is to close the edit view. Not only does making this a link incorrectly communicate its purpose—visually and otherwise—but it saddles the form “control” with lots of features, like bookmarking and middle-clicking, that have incorrect behavior. There are many plausible ways these buttons could be implemented, but none of those implementations should present themselves to the user as a link. With Button Actions, this entire UX could be implemented with just HTML. The first two buttons use existing HTML features, the second two buttons are made possible by Button Actions. (I’m also taking advantage of Triptych’s DELETE support , but you could do the URL method hack without it.) Philosophically, Button Actions create a generic control that can redraw the current context with a network request. Buttons already have the ability to do this with certain limitations; this proposal removes those limitations. Practically, this allows web authors to implement state transitions by navigating to views. Those views might even already exist as standalone destinations, in which case authors can trivially re-use existing routes while representing the action correctly in the UI. This is a great pattern that HTML should encourage! Unfortunately, without Button Actions, erroneously making this button a link is the only way that we have to implement this interface without scripting. This is obviously an anti-pattern, but it’s an anti-pattern supported by major design systems, because buttons lack the ability to do basic navigation without forms. When building a website that works without JavaScript ( a requirement for UK government sites ), links are the only choice. The US Web Design System (USWDS) even contains an official affordance for it: Add to a link and it will look like a button. Making a link look like a button, however, does not make the link behave like a button. USWDS uses JavaScript to implement spacebar activation , but JavaScript can’t do anything about the litany of other behaviors that differentiate buttons from links, like context menus. Links (even those with ) will still look like links in reader mode or other custom views. That’s the fundamental consequence of violating HTML semantics—the page will be broken for some users because authors cannot possibly account for all the different ways that people interact with a web page. The web simply wouldn’t work if they had to. Navigations are the broadest tool that web authors have to control the user experience—HTML just needs to complete the ’s ability to trigger them. Doing so makes the web simpler, safer, and more accessible for all. If you’d like to support the effort, the best way is to like the Button Actions issue on GitHub and share examples of why the proposal would be valuable to you. “Save” updates the comment with whatever is in the “Save Draft” saves the content of the without publishing it “Cancel” closes the editable form “Delete” removes the comment entirely Big shoutout to The Django Software Foundation for their support of this proposal ! I am currently working on an analysis to demonstrate that Button Actions do not introduce any new XSS vulnerabilities to existing web sites. Supporting this proposal doesn’t resolve that issue, but it does demonstrate to WHATWG that web authors have this need and that it’s worth studying. This blog focuses on buttons that trigger GET requests without forms, because that’s where the overlap with links is, but buttons that trigger unsafe requests without forms are also very useful. requests are probably the most common use-case, because they usually don’t require any additional data. One interesting case for buttons that trigger or requests without a form is “likes” on social sites . HackerNews , for instance, uses links for upvotes, which is in wild violation of HTTP semantics. I understand why they do it though: it’s simpler and works without JavaScript. That’s why it’s necessary to make Button Actions not just possible, but convenient. The proposal addresses all the existing workarounds for the lack of this functionality and explains why they’re not sufficient. The big picture goal with Triptych to is to give web authors a simple and semantic way to model a full CRUD lifecycle in HTML, because that’s all the vast majority of web services need to do. All the Triptych Proposals complement each other—Button Actions are even more useful with additional methods and partial page replacement —but I try to make the case for each one in isolation, both as an anti-logrolling mechanism and because they are genuinely useful on their own.

0 views
Lalit Maganti 2 weeks ago

How I Find Problems to Solve as a Staff Engineer

Note: this post was revised after publishing for increased clarity, based on reader feedback . “How do you find problems worth working on?” a senior engineer I mentor asked me recently. He’s trying to make the jump to staff engineer and realized that the role isn’t just about doing the work he’s assigned. He also needs to get involved in figuring out what his team and org should be building. Someone else had suggested blocking out time in his calendar to think about the bigger picture. He’d tried that, but hadn’t found it productive, so he asked if I had any alternatives. I told him I rarely find good problems by staring at a blank page and trying to “think strategically.” Instead, I act like a sponge. I listen to the stream of day-to-day noise, absorb the problems people are having and let them sit in the back of my mind. Over time, some fade away while connections begin to appear between others that initially seemed unrelated. Eventually, I start to see what’s really slowing people down and what my team or I can do about it. I’ve worked with many engineers who’ve never really tried this. They wait for managers or leads to identify opportunities, then demonstrate their value by solving the hardest assigned problems. That can absolutely lead to promotion. But the projects that have made the biggest impression in my career were the ones where I found and solved an important problem my leaders did not yet realize existed. One caveat: my experience comes mainly from working on infrastructure and developer tools at large companies, on teams where engineers have a lot of bottom-up autonomy to influence their roadmaps. In a more top-down environment, there may simply be less room to work this way. People love talking about the problems they are facing: in meetings, chat threads, presentations and email. They explain why their work is hard, complain about what slows them down and describe what they wish they could do. When something overlaps with my area, I start pulling on the thread. I might ask, “If X existed, would it solve your problem?” or point them at an existing feature in a product I own and ask how much of their use case it covers. Users often ask for a particular solution instead of explaining their root issue. Rather than taking the request at face value, I keep digging until I understand what they are trying to accomplish and why existing products do not work for them. As a natural introvert, this sort of ambient listening works particularly well for me. I don’t need to fill my calendar with speculative meetings just to find ideas; there is already an enormous amount of useful information flowing around me during a normal week. When a problem seems worth exploring, though, I become more active; I need to see how it affects the team’s day-to-day work. I’ll sit with them as they walk me through their workflows and the bugs they’re investigating. When I can, I’ll try working through some of those bugs myself. Seeing the problem firsthand makes it easier to separate what the team actually needs from the solution they asked for. I also seek out people who see more of the organization than I do: those who own critical systems, work across several teams or have particularly deep insight into the work downstream of my team. I’ll arrange a 1:1 or coffee chat and ask about interesting problems they’ve come across. They may have already seen the same issue in several places and started connecting the dots, giving me a head start on patterns I might otherwise have taken much longer to notice. Several times, I’ve been burned by moving too fast. I became excited by a request from a vocal team, built the feature and watched them barely use it. Their priorities had changed, or the request had come from a one-off investigation that no longer mattered. How eager a team was in that moment wasn’t the same as how important the feature was relative to everything else my product needed to support. By hyperfocusing on their request, I lost sight of the bigger picture. That taught me to let potential problems pile up. Listening the way I do leaves me with far more of them than I could possibly solve, and not all deserve action. Most don’t need to turn into projects the first time I hear about them; waiting can be a superpower. Waiting means the same problem might pop up independently in different teams, making it a higher priority to solve. Or problems that look different on the surface might turn out to have the same shape, so I can address several use cases in one shot. Or, as I’ve learned painfully, the requesting team didn’t even care that much in the first place. Instead, I make a mental note and revisit the problem if it comes up again. Other engineers I know write this sort of thing down more systematically. The mechanism is a personal choice: everyone has to figure out what works for them. What matters is keeping unresolved problems around long enough for more evidence to accumulate. Waiting helps me collect evidence, but that alone doesn’t tell me what to build. I still need to work out whether the problems I’ve retained are genuinely related and what, if anything, could address them together. Perfetto, the performance debugging tool I work on, is a good example. It displays recordings of system activity on a timeline made up of rows called “tracks.” Over a couple of years, teams kept asking for small, specific additions to the UI. One wanted a command to keep their preferred tracks pinned to the top of the screen; the next team wanted the same, but for a completely different set of tracks. Others wanted Perfetto to open already zoomed in on a particular part of a recording, or to show a custom aggregation tuned to what they cared about. A few had stopped waiting for us and built elaborate workarounds with bookmarklets. 1 By the time enough of these had piled up, my head was the usual tangle: the requests themselves, the constraints on each and a handful of half-formed solutions. I’ve learned not to force a solution by just sitting at a desk and thinking. Instead, my best untangling happens on long, aimless walks around London, where connections come more easily when I’m not trying to force them. What I eventually realized was that none of these teams really wanted the specific feature they’d asked for. Each wanted to personalize Perfetto for their own workflow without imposing their choices on everyone else. The underlying need wasn’t any one feature but rather the ability to extend the UI. When a connection like that finally clicks, it’s one of the best feelings in the job: several awkward requests collapse into a single idea, and possibilities open up that none of them hinted at on their own. That feeling, though, is exactly when I have to be careful, because a common shape is only a hypothesis and elegance is not evidence. When it happened with extending the UI it turned out to be real, but I’ve been fooled before. In another recent case I was convinced that building a transparent caching system for querying Perfetto traces would solve issues with sharing large traces and repeated queries. It was only as I wrote the RFC and built a prototype that I realized the elegance was a lie: the two problems wanted genuinely different solutions. I reluctantly split the design in two, both halves of which have since shipped. 2 You’d think this would be the moment I start building, but it usually isn’t. How far I go depends on how sure I am that the idea works and that people actually want it. If something is useful and low-risk enough, I act straight away: I send the change and let my manager know. When I’m unsure whether an idea will work or how much effort it will take, I build a throwaway prototype instead; it exposes the failure points and gives me something concrete for others to react to. And when an idea is big but I’m convinced by it, I commit to the full effort: weeks or months of work and the hard yards of building support across other engineers and teams. Through all of it, I’m not only trying to convince other people; I’m also trying to convince myself. Sometimes the honest answer is to stop: if people don’t see the value I do, or we hit a major technical wall, I’d rather drop the idea now than build something no one uses or that becomes a maintenance nightmare. And sometimes it holds up but the timing is wrong, so I park it, ready to spring into action the day it becomes an org priority. When an idea does hold up, I don’t necessarily need to be the person who builds it. I might implement it, someone else on my team might, or it might change what the org focuses on. Finding and shaping the right problem can have an impact even when I don’t own the implementation. The Perfetto extensions idea was worth that full effort. We were already building plugins to modularize the UI, but they weren’t enough: teams had to open source all their plugin code, which wasn’t an option for many internal use cases. So before building anything new, I took the problem and my proposal to my manager, teammates and the client teams. I ended up writing two RFCs, having several 1:1s and giving a couple of talks, refining it as the feedback came in. In the end, I designed and implemented macros as “lightweight extensions”: a way to automate actions in the UI without writing a plugin. Extension servers took the idea further by letting teams share their macros. Instead of implementing every requested feature ourselves, we gave teams ways to adapt Perfetto to their own needs. Dozens of teams inside Google now use macros and extension servers, and several other companies use extension servers internally too. The more often I go through this process, the easier it becomes. When I show genuine interest in someone’s problem, ask useful questions or help solve it, they remember. They start coming to me earlier and bring me into conversations with other people facing related issues. That gives me a wider view of what is happening across the organization, making it easier to spot patterns and build things people actually need. Solving one of those problems brings me into more conversations, and the loop continues. Those successes build the kind of trust that comes from long-term stewardship . Early on, I had to turn many of these ideas into something real myself to prove that my judgment was sound. Over time, my manager and org gave more weight to my assessment of what mattered. That allowed me to influence the roadmap without needing to own every project. This differs from the idea that becoming a staff engineer means replacing technical work with meetings and coordination. For me, conversations are inputs into what I build, not the end result. That is what I wanted my mentee to understand: finding problems worth solving isn’t separate from the rest of the job. It comes from staying engaged with people’s work long enough to see what no single request can show you. These workarounds used bookmarklets to run JavaScript against Perfetto’s internal UI APIs.  ↩︎ The original proposal was to use a transparent cache for repeated queries and faster reopening of large traces. As I worked through it, I realized repeated queries were better served by keeping sessions warm in memory, whereas reopening was better served by explicitly exporting a trace into a format designed to load quickly. A transparent disk cache could also retain multi-gigabyte files without the user realizing and would need a new system to manage their lifetime. The proposal was ultimately replaced by warm sessions and streaming table export .  ↩︎ These workarounds used bookmarklets to run JavaScript against Perfetto’s internal UI APIs.  ↩︎ The original proposal was to use a transparent cache for repeated queries and faster reopening of large traces. As I worked through it, I realized repeated queries were better served by keeping sessions warm in memory, whereas reopening was better served by explicitly exporting a trace into a format designed to load quickly. A transparent disk cache could also retain multi-gigabyte files without the user realizing and would need a new system to manage their lifetime. The proposal was ultimately replaced by warm sessions and streaming table export .  ↩︎

0 views
Ahmad Alfy 2 weeks ago

Testing Google’s “modern-web-guidance” skill against a real React app

LLM-assisted frontend work has a particular failure mode. The model confidently writes code that was best-practice in 2021. It reaches for , hand-rolls a dark-mode toggle with a class on , or disables the submit button to “prevent” invalid input. None of it is wrong exactly. It’s just a few years stale, because the training data is a few years stale and the web platform moves faster than that. Google Chrome’s skill is a direct attempt to fix that. It’s not a linter and it’s not a codegen tool. It’s a search index over a curated set of best-practice guides , meant to be consulted before you write HTML/CSS/client-side JS, so the pattern you reach for is the current one. I wanted to know whether it actually earns its place in the loop. So I pointed it at a real codebase, the React frontend of a project-assessment internal tool I’ve been building, and treated it as an auditor. This is what came back. There’s no magic. It’s two commands over : returns a ranked JSON list. Each hit has an , a , the web , a , and a semantic score. returns the guide as markdown. That’s the whole interface. The intelligence is in (a) the quality of the guides themselves and (b) whether the semantic search puts the right guide in front of you. Everything below is a test of both. The app is a Vite + React 18 questionnaire. You answer about 10 questions, it computes a recommended tech stack client-side, and you can save, label, and annotate assessments. It runs to 42 source files. What matters for this exercise is that it’s form-and-input heavy but has no images and no marketing-page concerns. So the relevant guidance is going to be about forms, inputs, theming, and layout, not LCP hero images. I did a quick inventory first. The tells were immediate: Then I let the skill tell me what to do about each. I searched for . The top hit came back at 0.75 similarity , the highest of the whole session: The app’s current theming is a wall of light-mode hex: The retrieved guide is refreshingly opinionated about what’s mandatory versus optional. The two non-negotiables: That single declaration is the highest-leverage line the audit surfaced. Without it, even a perfectly hand-themed dark palette leaves the native scrollbars, widgets, and the initial paint canvas stuck in light mode. That’s the exact “white flash on load” that makes a dark site feel broken. Beyond the mandatory two lines, the guide shows how to define color tokens with , so each token carries its light and dark value in one place. Applied to this app’s theme file, the change is small. Every hardcoded hex becomes a pair, plus the two mandatory declarations: And updating the is just one line: (The dark values are illustrative inversions. The point is the shape of the change, not the exact palette.) What surprised me is that the guide doesn’t stop at CSS. It carries a section on the design of a theme toggle. This is part of the guide’s own text. You can read it with , or straight on GitHub in the dark-mode guide . Its UX considerations subsection makes the sharpest call, arguing that you should not build the toggle most of us reflexively build: DON’T expose all three states (system, light, dark). … Two of the three options always produce the same visual result, violating the principle of feedback. Instead it argues for a two-state control, “follow the system” and “the opposite of the system.” It also spells out the edge case that trips people up. If a user pins dark and then switches their OS to dark too, the site must stay dark rather than flip. That’s product judgment sitting inside a CSS guide, and it’s exactly the kind of thing a model won’t reliably volunteer on its own. Finally, because is newer than , the guide hands over the fallback so you don’t have to reason it out. You degrade through , then upgrade with where exists. It also ships a copy-paste script to prevent the theme flash for users who have pinned a non-default choice. That script is a plain inline one, deliberately not and not a module, so it reads the saved preference before first paint. And that brings up the skill’s best structural feature. When a guide leans on anything newer than the long-settled web, it keys its browser-support advice to Baseline . For , the dark-mode guide returned that it’s widely available and has been Baseline since 2022-02-03. For , it returned that it’s newly available and Baseline since 2024-05-13. This matters because it turns “should I use this?” from a vibe into a decision rule. The skill’s own instructions say Baseline-Widely-available features are safe to use unfenced, while newer features must carry the fallback the guide provides, unless you’ve declared a custom browser-support policy. In other words it defaults to safe, and it tells you exactly where the risk line is instead of leaving you to guess. is safe to just ship, while gets a -guarded fallback. That’s the correct call, and it made it without me having to ask. I searched for . That surfaced the guide at 0.50 similarity, with and an guide right behind it. Here’s the app’s save surface, lightly trimmed: The guide’s very first rule is blunt about it. “DO use the element to wrap interactive controls… DON’T use for primary submission buttons.” The rename field and the notes editor elsewhere in the app repeat the same -plus- shape. The practical cost of the current approach isn’t abstract. Because there’s no , pressing Enter in the label field does nothing , and that’s a reflex every keyboard user has. The fix is small, and the guide hands it over directly, including the AJAX-friendly submit handler: Wrap the input and button in a , make the button , and Enter-to-submit comes back for free, along with native form semantics for assistive tech. I’ll give the skill credit for a fair grade, too. One thing the app already does right showed up in the same guide. The save button disables itself while a save is in flight, and the guide explicitly blesses that. “DO disable the button after a valid submission is clicked to prevent double-posts.” This is the opposite of the anti-pattern from the intro. Disabling after a valid click to stop double-submits is good, while disabling up front to block an incomplete form is the dead end. A good auditor tells you what to keep, not only what to change. I searched for . It surfaced a cluster of tightly-scoped guides, , , and , all built around and . This is where the guides go deeper than a model’s default answer. Ask a chatbot “how do I validate a form field” and you’ll usually get an handler that yells the moment you type one character. The guide instead ships a timing matrix : The guide even boils it down to a single rule. “Validate on to avoid premature warnings while typing, and reset error states on as soon as the user attempts a correction.” The modern platform gives you this essentially for free via the pseudo-class, which only matches after the user has interacted. The app’s ad-hoc error paragraphs are accessible, which is another thing it got right, but they’re wired by hand where the platform now has a purpose-built primitive. The guide’s section 3 covers , , and , the attributes that tune autofill and the on-screen keyboard. None of the app’s inputs use them, and this is the one finding where the honest answer is a polite no. The questionnaire is almost entirely , where you pick one of a handful of options. Radios don’t take an or an token, because there’s no keyboard to optimise and nothing to autofill. The only free-text fields in the whole app are a “label” and a “notes” box, and neither maps to a standard autofill value. So the guidance is correct in general and largely irrelevant here, and noticing that is the actual work. The tool returns a rule. Deciding it doesn’t apply to a radio-driven form is a judgment call it can’t make for you. This is the clearest example in the whole audit of why the skill is only half the loop. One line from the section does still land universally, though. Text inputs should be or larger, because anything smaller triggers an auto-zoom on iOS Safari the moment the field is focused. The app has in two files. The well-known modern fix is (dynamic viewport height), which accounts for mobile browser chrome that ignores. On iOS Safari, is measured against the expanded viewport, so the bottom of a layout sits behind the address bar. My query returned the broad and guides rather than a laser-focused “use dvh” atom. The right answer is almost certainly inside those guides, but the search didn’t hand me a -titled hit the way it did for . Which is a fair segue into the honest assessment. is not going to catch your bugs and it won’t rewrite your components. What it does is remove the single most common source of stale frontend code, the confident-but-outdated pattern. In one afternoon pointed at a real app, it correctly flagged a missing declaration, a set of forms that skip native submission, a validation approach that predates , and a pile of missing input attributes. For each one it handed over current, Baseline-checked, copy-pasteable guidance, while also telling me which of my existing choices to leave alone. One reframe stuck with me. It’s less a tool you run and more a standard you consult . The best time to reach for it isn’t during a cleanup audit like this one. It’s the moment before you write a component, when the model in the loop (human or AI) is about to reach for the pattern it already knows. Half the time, the pattern it knows is three years old. This is the cheap check that catches it. Zero elements. Every data-entry surface is a bare plus a . No , , or on any input. A hardcoded light theme. defines tokens as literal hex values, with no , no , no dark variant. in two places. Some genuinely good instincts too, like / grouping, on errors, and . The guides are high quality. They read less like scraped blog posts and more like a curated reference assembled by people who live in the web platform, close to the specs and the browser internals, but writing for the developer who actually has to ship. The mandatory/optional split, the timing matrices, and the “don’t build a three-way toggle” UX arguments all read as earned judgment, not a spec dump. Baseline-keyed fallbacks where a feature needs one. This is the single best thing about it. It converts “is this safe?” into a date comparison and provides the exact fallback when the answer is “not yet.” You don’t have to look for the fallback, it comes with the guidance. It grades fairly. In two places it validated code the app already had right. An auditor you can trust to say “keep this” is one you’ll actually keep running. Framework-agnostic by design. Every guide is HTML/CSS/DOM, and adapting the pattern to React was trivial. Nothing assumed a framework, so nothing fought mine. It’s local, self-contained, and keyless. The semantic search runs on your own machine through a small on-device model, so the matching itself makes no network calls and there are no API keys to manage. The npm package ships with no extra dependencies, which keeps latency low and the supply-chain surface small, and the CLI can run fully offline. By default the tool reports anonymous usage statistics to Google, including your search queries and guide retrievals, which you can turn off by setting . It doesn’t read your code. You (or your agent) do. This is the big one, and it’s worth being exact about, because it changes how you run the skill. Nothing in this audit was automatic. The app had to be read, the suspect patterns spotted, each one turned into a search phrase, and the returned guidance compared back against the actual lines. There are two ways to do that. You can drive it by hand, deciding what to search, reading the guides, and applying them yourself. Or you can hand the whole loop to a coding agent, which is what I did here. The agent inventoried the frontend, chose the queries, retrieved the guides, and did the comparison, while the skill only ever answered “here is the current best practice for X .” Either way, the skill supplies the standard and something else supplies the code-reading. Point it at a codebase with no idea what you’re looking for and it hands you nothing back. Semantic search has a recall ceiling. hit at 0.75, but the answer never surfaced as its own result. When a query returns only broad category guides, you have to retrieve a large omnibus guide and read it yourself, which brings up cost. The guides aren’t small, but the skill is upfront about it. Every search result carries a in its JSON. The guide reports about 4,500, and about 7,100. I didn’t measure those myself, because the tool hands them to you before you fetch, so you can weigh the cost. Retrieving a few of them still meaningfully fills a context window. That’s fine for a deliberate audit, but something to watch if you wire it into every edit.

0 views
マリウス 2 weeks ago

I Regret Migrating to Codeberg

My primary reason for leaving GitHub was not about a single feature or a single outage, but about the “enshittification” of the platform under Microsoft ’s ownership. The web interface got rewritten into a sluggish pile of JavaScript that either broke things which used to just work, or made them so horribly slow that using them became a PITA . Beyond the technical decay GitHub had turned into de facto “public infrastructure” in much the same way that WhatsApp has , hosting the source code of a very large share of the world’s software and, through that, giving Microsoft a degree of leverage and surveillance over everyone’s projects, and by extension everyone’s digital lives, that no single company should hold. On top of that, stories about legitimate developers losing their accounts due to arbitrary bans by Microsoft only reinforced the feeling that it would be a good idea to at least have a backup somewhere else . Codeberg looked like a viable alternative. It offered free and open-source projects a reputable home and, more importantly, an equally free one, run by a non-profit association rather than a subsidiary of the largest software vendor on the planet. Unfortunately, the latest update to its terms of service seems to mark a first step in changing one part I moved there for, namely the “freedom” part. Every project I’ve published so far was built with 100% human stupidity rather than “artificial intelligence” , or, more accurately, LLMs . I don’t hold particularly strong feelings about Codeberg banning projects that are predominantly LLM -driven, at least not feelings as strong as the ones I hold about the simultaneous ban of legitimate cryptocurrency projects, which reads as though it got lumped in for no reason other than that most people still remember the villain-du-jour that crypto was in the years before LLMs took that title. The two clauses landed within days of each other, the LLM prohibition on the 29th of June and the cryptocurrency prohibition on the 2nd of July, both as Assembly 2026 proposals, and the terms now file the latter under, of all things, “content that harms the reputation of Codeberg” , which sounds like legalese for “we don’t have a solid reason or an actual number of bad precedents to categorically ban it” . The announcement blog post , however, reads very poorly, and the section titled “The development team of none” is the worst of it. It states: Using LLMs to work with your code gives you a kick of adrenaline. You can develop at a rapid pace, build things as if you had a large team. Only that you have none. In fact, you are (often) alone, working with a statistical machine that turns energy into code. And, a little further down, it says: It seems like many ‘vibe coders’ don’t realize that they don’t actually have a community around them. This is out of touch with how most free software gets made. The majority of FOSS developers are one-man-shows, and the only cOmMuNiTy they have around them are the users requesting features or reporting bugs while most of the time not contributing in any form whatsoever. I’ve been publishing silly little tools for decades, predating this website and even GitHub itself (remember when SourceForge was the hot sh.t ?), and not one of them has ever had an actual “community” around it, at least not in the romanticized sense that Codeberg paints in that post. I’m a lone wolf who codes everything by hand and spends an absurd amount of time doing exactly that, and the notion that an LLM is the thing separating a real project with a real community from a fake one does not hold up once you look at how the average useful little tool on any forge comes to exist in the first place. It’s frankly a bit snotty of Codeberg to make this argument at all, considering that the platform effectively lives inside the Forgejo bubble, and Forgejo mutinied inherited its community of active contributors from Gitea , who had spent the better part of six years building that community before Forgejo even existed. A project that acquired its own community by hard-forking someone else’s, then turned around to lecture solo developers about not having one, is a difficult position to argue from with a straight face. In addition, Codeberg conflates “having a community” with “being legitimate software worth hosting” , when the bar for a personal project has always been a working build, ideally a license, and maybe a README, and not a channel full of contributors. A good deal of what makes the small, single-author tool ecosystem worth having is precisely that it doesn’t need a community to justify its existence, and a forge whose entire selling point is hosting the code of individuals is an odd place to argue the opposite. The part that bothers me isn’t the specific ban on LLM projects, or the specific ban on cryptocurrency projects. It’s that a hub built around “free software” is now telling its users which kinds of software are deemed good and which are not, and that is closer to censorship than it might seem. Once a platform writes into its terms that an entire category “harms its reputation” and can be removed on that basis, the deciding factor stops being whether the code is legal, or functional, or useful, and becomes whether it aligns with a position the platform has taken. I would argue that a significant share of the projects caught by a blanket ban of that kind are legitimate software rather than vibe-coded slop or sh.tcoin implementations. Every platform I can think of that took this approach became divisive the moment it started enforcing an ideology on its users, whatever that ideology happened to be, and however justified it looked at the time. The mechanism is always the same, where a real problem shows up, an unpopular category becomes the obvious culprit, the platform bans the category instead of addressing the problem, and that ban then becomes the precedent for the next category, and the one after that. The category that is uncontroversial to ban today is the reason the mechanism exists tomorrow, and the users who applauded the first ban rarely get asked about the second one. I do acknowledge that both categories aren’t free of problems. LLM -driven repositories do strain infrastructure, do generate unmanageable volumes of low-quality issues and pull requests, and do carry real questions about copyright and code provenance, all of which Codeberg names in its post. The cryptocurrency space, in turn, might have produced more outright scams than almost any other corner of software. However, a categoric ban on the villain-du-jour is not a solution to any of that. We now even have people like Linus Torvalds making the fairly reasonable argument that an LLM is just a tool , and “clearly a useful one” , with a legitimate place in Linux kernel development when it’s used carefully and its output is held to the same standard as everything else. If the maintainer of the largest and most consequential open-source project on the planet can treat LLMs as a tool to be judged on its results rather than a category to be banned on sight, a backyard code forge can manage the same. I, too, am worried about the impact of LLMs on tech, and on society in general, going forward, and I’d guess I’m about as worried as whoever wrote Codeberg ’s policy. I just don’t believe that banning content, which is very much what this amounts to, is the way forward. What I wish Codeberg had reached for is a solution that treats the actual problem, which by their own account in that same post is resource consumption and the infrastructure cost that comes with it, as an actual resource problem. A change to the terms of service could have required authors to tick a checkbox declaring that a repository contains LLM -generated code, or is cryptocurrency-related, and those repositories could then be segmented onto a separate tier of infrastructure that doesn’t get the same resources as everyone else. A tier that carries specific quotas, and that might require the author to pay for what they consume. Declaring the truth honestly would (at least at first) cost nothing, and failing to declare it, then getting caught, could be met with exactly the permanent, immediate ban that Codeberg is now applying to entire categories from the outset. Similarly, projects that carry the LLM or Crypto label could carry automatically displayed disclaimers that explicitly state that Codeberg is in no way responsible for the quality or correctness of this specific repository. Heck, they might even go as far as to blatantly state that Codeberg does not approve of the use of LLMs or Cryptocurrencies in those warnings, to make extra-extra-extra sure that people get it and that there is no “reputational risk” for Codeberg . An approach like this puts the cost of resource-hungry projects onto the people creating them, and it keeps the shared resources for the projects that were the reason the platform exists. All of that without Codeberg having to decide which categories of software are ideologically acceptable in the first place. The “we ban everything upfront that we don’t agree with” approach is the wrong signal to send, and it is a very slippery slope. Despite not owning a single project that falls into either banned category, I’m now going to look into setting up my own public Git host, and I’ll move off Codeberg only a few months after moving there , because of this. Not because of the bans themselves, but because I don’t want to depend on a platform that rewrites its terms of service on a whim, without properly announcing that the change was even under consideration, and without giving its users a way to weigh in. The decisions did go through Codeberg ’s own Assembly 2026 , which is more process than most platforms bother with, and yet as an ordinary user I found out about it the way probably most people else did, through a dark blue banner at the top of the site on the day it was already settled. While I appreciate the info about the ToS change, I wish I’d gotten a banner back when the platform was still deciding whether to go down this road, and I wish it had linked to a discussion thread, or at the very least a poll, so that I could have voiced the concern I have, which is about the freedom of the platform as a whole, rather than about any single category that ended up banned.

0 views
Simon Willison 2 weeks ago

OpenAI’s accidental cyberattack against Hugging Face is science fiction that happened

This story is wild. The short version: OpenAI were running a cybersecurity test against an unreleased model, with the model's guardrail features turned off. Rather than solve the test, the model broke its way out of OpenAI's sandbox, then found exploits to break in to Hugging Face, all so it could cheat on the test by stealing the answers. Along the way it helped make the strongest case yet for how the imbalance of model availability is hurting our ability to secure our software. We currently have three documents to help us understand what happened here. I hadn't seen the ExploitGym paper before and it's a really interesting one. Authors from UC Berkeley, the Max Planck Institute, UC Santa Barbara, and Arizona State designed a new benchmark for evaluating models on their ability to turn a reported vulnerability into a concrete exploit. OpenAI, Anthropic, and Google provided feedback and helped run the benchmark against their models. The benchmark "comprises 898 instances derived from real-world vulnerabilities that affected popular software projects" - including the Linux kernel and V8 JavaScript engine. Here's the paragraph that best represents their benchmark results: Among all configurations, Claude Mythos Preview and GPT-5.5 achieve the highest success counts (157 and 120 successes, respectively), demonstrating that current frontier agents can exploit a substantial subset of real-world vulnerabilities under controlled conditions. GPT-5.4 also solves a notable 54 tasks, placing it in an intermediate tier. The remaining model–agent pairings solve fewer than 15 tasks each, underscoring that end-to-end exploitation remains challenging and sharply differentiates today’s frontier systems. Notably, Claude Opus 4.7 achieves fewer successes than Claude Opus 4.6 despite being a newer checkpoint, and does so at substantially lower cost on the full set. Trace inspection reveals that Claude Opus 4.7 and Gemini 3.1 Pro frequently conclude early after judging the target vulnerability non-exploitable. The paper also describes the approach they took to preventing the agents from cheating by going outside the parameters of the test. This becomes relevant in a moment! Outbound connections are restricted to a curated allowlist that permits routine package installation (Ubuntu apt repositories and PyPI) and fetching the toolchains required for building V8. All other external endpoints are blocked. The paper concludes with this (emphasis mine): Our results show that autonomous exploit development by frontier AI agents is no longer a hypothetical capability . While current agents are not yet reliable across all targets, they already exploit a non-trivial fraction of real-world vulnerabilities , including complex targets such as kernel components. This rapid emergence is itself a central finding, showing that capabilities that would have seemed implausible are now present in deployed frontier models. An important detail here: this paper isn't about discovering vulnerabilities; it's about being able to take those vulnerabilities and turn them into working exploits. When Anthropic first restricted access to Mythos back in April they talked about this capability as well. A model that can act on vulnerabilities is a lot more dangerous than one that can just discover them. One of the ways Fable differs from Mythos is that it's more likely to refuse to weaponize vulnerabilities in this way. I get the impression the US government did not understand that distinction when they banned Fable last month . The first hint we got of the attack was in this blog post by Hugging Face on 16th July 2026: A malicious dataset abused two code-execution paths in our dataset processing (a remote-code dataset loader and a template-injection in a dataset configuration) to run code on a processing worker. From there, the actor escalated to node-level access, harvested cloud and cluster credentials, and moved laterally into several internal clusters over a weekend. I hope they release more details about the code that pulled this off. I'm assuming this means packages using the datasets library , a Hugging Face project for bundling up and sharing datasets on their platform. That library used to execute arbitrary code but has been steadily locked down over time, with the 4.0.0 release in July 2025 removing the flag entirely. Assuming the attack used that library it must have either abused pickle serialization in some way, found some other non-obvious code execution path, or (most likely) specified as the dependency. The campaign was run by an autonomous agent framework (appearing to be built on an agentic security-research harness - used LLM still not known) executing many thousands of individual actions across a swarm of short-lived sandboxes, with self-migrating command-and-control staged on public services. This was a sophisticated attack! Then Hugging Face hit a wall: they tried to use "frontier models behind commercial APIs" - I'm guessing from Anthropic and OpenAI - to help analyze the attack, and were blocked: When we started the log analysis, we first used frontier models behind commercial APIs. This did not work: the analysis requires submitting large volumes of real attack commands, exploit payloads, and C2 artifacts, and these requests were blocked by the providers' safety guardrails, which cannot distinguish an incident responder from an attacker. They switched to their own self-hosted instance of MIT licensed GLM-5.2 and it helped them figure out what was going on. This indicated a fundamental asymmetry between the defending team and the (so-far unknown) attacker: We do not know which model powered the attacker's agents, whether a jailbroken hosted model or an unrestricted open-weight one; either way, the attacker was bound by no usage policy, while our own forensic work was blocked by the guardrails of the hosted models we first tried. As a useful indicator of how seriously they took the attack: [...] Finally, we have also reported this incident to law enforcement agencies. So who was responsible for this "autonomous agent framework"? It turned out to be OpenAI themselves. Five days later, on July 21st , OpenAI revealed the culprit. They had been running the ExploitGym benchmark against a new, as-yet undisclosed model, and that model had been operating way outside its intended parameters (emphasis mine): After investigating, we now know that this particular incident was driven by a combination of OpenAI models — including GPT‑5.6 Sol and an even more capable pre-release model, all with reduced cyber refusals for evaluation purposes — while being internally tested on a benchmark⁠ [ExploitGym] of cyber capabilities. [...] We estimate maximal cyber capabilities by running this evaluation without production classifiers used to prevent models from pursuing high-risk cyber activity . Our benchmarks run in a highly isolated environment, with network access constrained to the ability to install packages through an internally hosted third-party software that acts as a proxy and cache for package registries. The models identified and chained vulnerabilities across OpenAI’s research environment and Hugging Face’s production infrastructure to obtain test solutions directly from Hugging Face’s production database. All evidence suggests that the models were hyperfocused on finding a solution for ExploitGym, going to extreme lengths to achieve a rather narrow testing goal. It's pretty clear what happened here. OpenAI removed safety filters for an in-progress model, locked it up in a sandbox and told it to solve the ExploitGym problems. Given the absence of guardrails there was nothing to prevent the model from attempting to break out of that sandbox, break into Hugging Face, and read the answers from there instead. OpenAI's sandbox used the now common pattern of allowing access to an allow-list of sites related to installing additional packages via a controlled proxy. In OpenAI's words: While operating in our sandboxed testing environment, our models spent a substantial amount of inference compute finding a way to obtain open Internet access , in pursuit of solving the evaluation problem. To gain access, the models identified and exploited a zero-day vulnerability (which we’ve now responsibly disclosed to the vendor) in the package registry cache proxy . With this access, our models performed a series of privilege escalation and lateral movement actions in our research testing environment until the models reached a node with Internet access. So step one was to break out onto the public internet. The model then broke into Hugging Face to find the answers: After gaining Internet access, the models inferred that Hugging Face potentially hosted models, datasets and solutions for ExploitGym. Knowing this, the model searched for and successfully found ways to gain access to secret information that it could use to cheat the evaluation. In one example, the model chained together multiple attack vectors, including using stolen credentials and zero-day vulnerabilities to find a remote code execution path on the Hugging Face servers. Chaining together multiple attack vectors is exactly the kind of thing these new models can do, where previous generations of models might have failed. I wrote last month about how Claude Fable is relentlessly proactive , when I noticed it spinning up custom web servers and deploying CORS tricks on my own laptop just to help debug a WebKit CSS issue. It turns out relentless proactivity is the defining trait of this new generation of Mythos-class models. If you set them a goal and give them a way to get there, even inadvertently, they will figure it out . There will inevitably be some people who dismiss this story as a dishonest marketing trick by OpenAI to make their models sound terrifyingly effective. I found 81 instances of the term "marketing" in the Hacker News discussion of the incident. To those people I say pull your heads out of the sand - you're now including Hugging Face in your conspiracy theories, just so you can deny the crescendo of evidence here! The best models we have today have the ability to both find and exploit new vulnerabilities. The ExploitGym paper itself concludes that "autonomous exploit development by frontier AI agents is no longer a hypothetical capability", and this incident is a perfect example of exactly that. One of the most infuriating details of this story is how Hugging Face, faced with an accidental and aggressive attack from one of OpenAI's models, were unable to then turn to OpenAI's models to help them fend off the attack. The frontier models we have access to are increasingly being constrained in how much they can help us protect our software, heavily influenced by the US government's ongoing threat of export controls. Claude Fable 5 wouldn't even proofread this article for me! It insisted on downgrading me to a less capable model. Meanwhile open weight models from China such as GLM-5.2, Kimi 3 and the new Qwen 3.8 Max appear to have none of these restrictions - and any restrictions that do exist can likely be fine-tuned out of them by modifying the weights These constraints are meant to make us safer. I think there's a risk that they are having the opposite effect. You are only seeing the long-form articles from my blog. Subscribe to /atom/everything/ to get all of my posts, or take a look at my other subscription options . ExploitGym: Can AI Agents Turn Security Vulnerabilities into Real Attacks? is a paper published on 11th May 2026 describing ExploitGym, a new eval suite for LLM-powered agent systems. Security incident disclosure — July 2026 by Hugging Face on 16th July 2026 describes how they detected an attack from an "agentic security-research harness - used LLM still not known" that breached some of their systems. OpenAI and Hugging Face partner to address security incident during model evaluation from OpenAI on 21st July 2026 confesses that it was their agent harness that did this, and that they're working with Hugging Face to clean up the mess.

1 views
Julia Evans 2 weeks ago

Some more things about Django I've been enjoying

Hello! I’m on a funny journey right now where I’m trying to learn how to make websites in a sort of 2010 style, where I have an SQL database and render some HTML on the backend. It’s kind of an interesting journey because it doesn’t necessarily feel “easy” to me to make websites in this way: I never learned how to do it in the 2000s or 2010s, and there’s a lot I need to learn. So here are some Django features that make building this kind of site feel more achievable than when I was trying and failing to use Go’s standard library or Flask. And I’ll talk about a couple of issues with Django I’ve run into. Previously the toolkit I felt confident with for making websites was: I really liked this frontend-heavy approach for these super simple applications but when I started thinking about making something with a lot of different pages (instead of literally just one page), I didn’t feel so excited about the options I saw that involved a lot of frontend code. So I figured I’d try the backend. Writing a backend-focused site that uses as little JS as possible feels the same to me in a way as writing a single-page JS website that does as little on the backend as possible, even though they might seem like opposites. In both cases I’m just trying to keep as much of the logic as possible in one place. Now for some thoughts about Django! I learned that I can define a “query set” class in Django with a bunch of methods with different statements I might want to use while constructing a query: Here’s how I use it in my view code once I’ve defined what all the methods mean: and here’s how I define the methods: The syntax for defining the filters isn’t my favourite, but I spend most of my time just using the methods, and it feels super readable and nice to use, and it makes me want to look into other query builder libraries in the future. In the past I thought “I know SQL, who needs a query builder?”, but this kind of structure does make it really nice to read. I found an example of someone who wrote their own small query builder in Python that I want to read later to think about whether I would enjoy using a more minimal version of this. There are a bunch of little quality of life filters available in Django templates that are super useful for generating HTML. The ones I’ve used so far are: These are all small things individually but I feel like it makes a big difference somehow to just have them available. I think my favourite template filter is : in this site sometimes we use filters like to decide what’s displayed. that will make a link to the same query string with one change, like this to link to the previous date: Or to remove the parameter: I still really love Django’s automatic database system. It’s amazing to be able to just edit a model to add a new field or whatever, and then Django automatically generates the migration. So far we have done 19 database migrations and I think there will probably be more! It makes a huge difference for me to be able to just easily change the database as my understanding of the problem changes. Django’s documentation sometimes offers the option of using class-based views and inheritance to organize the code in your views. For example I have four views that share a lot of code, and I could use inheritance to manage that by defining some kind of parent class and then having my other views inherit from it. I tried it out and I did not enjoy the experience of using inheritance to share code between views. I switched to using functions instead, sort of how this post advocates, and that was a lot more straightforward. I’ve never had a good experience using inheritance in Python and I don’t think I’ll try to use it again. But I don’t mind using inheritance to use the interfaces Django itself provides: for example if I want to define a query set I need to write something like . I don’t think too hard about it and it seems to work. (as a meta comment: I’ve been working on talking about my programming opinions by just saying “THING does not feel good to me, I prefer OTHER THING instead”. That post I linked to says that function-based views are the “right way”. I’m not very invested in whether it’s “right”, but it’s validating to know that other people feel similarly to me about inheritance) At some point the LLM scrapers discovered our site, and started sending us maybe 10 requests per second. I blocked them which is working for now, but it made me think about what the site’s capacity is. I’m used to writing Go backends where the performance situation is pretty straightforward (usually everything is just fast enough), and a Django site is very different. Some light load testing (with ( ) shows that right now we can serve about 2-3 requests per second (on a ~$10/month VM). It’s tempting for me to go down a rabbit hole where I do a bunch of profiling to figure out what’s slow and try to make it faster (there’s py-spy for that, and py-spy is great and super easy to use, and profiling is fun!) But I really don’t understand what I should expect in terms of performance from a Django site and how I should be thinking about at a higher level. Some things I haven’t figured out yet: I think one thing I’m learning about Django is that because it’s a Framework (tm), it’s easy to accidentally misconfigure it. For example, when I was thinking about why my site was slow just now, I read the django performance docs and I noticed a comment saying: Enabling the cached template loader often improves performance drastically, as it avoids compiling each template every time it needs to be rendered. When I’d done CPU profiling I’d noticed that it was spending a lot of time rendering templates! Maybe this could help me! Clicking through the link, I saw that the cached template loader was supposed to be on by default, but I’d turned it off by accident while trying to do something else. I think this “I turned off the cached template loader by default” things is an example of how I still find the django settings file to be pretty confusing and difficult. I guess I should just be careful when I go in there. After turning on template caching, it seems like the site can now pretty easily handle 12 requests per second or so without using all of the CPU. I have not carefully benchmarked the before and after but it seems like it’s made a pretty big difference. One thing that’s been surprising to me about Django performance is that I’ve always heard the advice “if you have a performance problem, check your database queries! Maybe add an index!”. But I’ve been running into a variety of performance issues (like this template caching thing) that are not because of slow queries, so instead it’s been more useful for me so far to start by running a CPU profile. And since I’m using SQLite, any slow database query problem will show up on the CPU profile anyway. Anyway I don’t want to get too far into site performance. Like I said it’s easy for me to get interested in profiling, but actually I know a lot about profiling and it’s not the most important thing for me to learn about. I might say more about what I’m enjoying (or having a hard time with!) about Django later. Trying to write some shorter blog posts recently. static site generators (like for this blog) static sites that do some fun stuff with Javascript (like this sql playground ) simple Vue.js single page apps with either a Lambda as a backend or a Go backend (like mess with dns ) translating plain text URLs into links, or line breaks into ( ) formatting dates ( ) , which takes a Python dictionary and automatically converts it to JSON and inserts it into the HTML as a tag in a safe way If I have a site that’s going to be getting occasional bursts of traffic, do I want to be able to scale up? Do I want to design the site so that more things can be cached? (and do I really have to? caches are so annoying to get right!) The django performance docs say that Jinja is faster for templating, do I want to think about switching templating systems? Those docs also say “{% block %} is faster than using {% include %}”, I wonder if it’s a big difference and if so why

0 views
Maurycy 3 weeks ago

Regressive JPEGs:

One of the cool features of JPEG files is that there's the option to save low frequency components first. This means that a partially downloaded image will be displayed at low resolution instead of being cut off. In the file, this works by breaking up the compressed data into multiple "scans", each prefixed with a header. Here's the first scan of a representive image: ... this one includes the lowest (DC) Fourier bin for all three color channels. The three color channels are YCbCr instead of the usual RGB. The luminance (Y) seperated because it must be high quality, but the color can be fudged quite a bit while looking fine. Very roughly: Y = G, Cb = B - G, Cr = R - G After it, the file contains eight more scans to fill in the rest of the data: Scan number Channels DCT bin range Precision 0 Y Cb Cr 0 - 0 Half (-1 bit) 1 Y 1 - 5 Quarter (-2 bits) 2 Cb 1 - 63 Half 3 Cr 1 - 63 Half 4 Y 6 - 63 Quarter 5 Y 1 - 63 Half 6 Y Cr Cb 0 - 0 Full 7 Cr 1 - 63 Full 8 Cb 1 - 63 Full 9 Y 1 - 63 Full Scan #0 contains a very low resolution preview of the image. Scan #1 adds some details to the luminance. Scans number two through five contain full low precision data. Scan 4 has an unusual spectral range because it's filling in the gap left by #1. That way, number 5 has full quarter precision data to build on. Scans six through nine add the final missing bit to bring the image to full quality. Given what I said about color being less important, it might seem weird that my example has the color data first: This works because the the chrominance is saved at half resolution (quarter pixel count). As a result, full chrominance data (Cr + Cb) only weighs half as much as luminance. Since each scan explicitly sets its spectral range , it should be possible to construct a JPEG file where future scans overwrite already rendered image data. Actually, it's very easy to do this: Concatenate multiple images with the same resolution and filter out the start-of-image, start-of-frame and end-of-image markers. This can be done in a hex editor, but I used a quick and dirty C program. When served over a slow network , this concatenated file will switch between multiple images: Click to open in new tab But, most decoders will give up after some number of scans : I think this is done to avoid a zip bomb style problem... but it prevents this from working on more than 9 frames, which is not enough for a proper animation. To do that, I'd have to minimize the number of scans in each frame. The simplest idea is to start with baseline JPEGs that only have a single scan. ... but it doesn't work: In progressive mode, a scan can't contain both AC (bins above 0) and DC (bin 0) data at the same time. This limitation doesn't exist for baseline mode, but the baseline decoder stops after the first scan. Since AC data must follow DC data, the smallest possible "progressive" JPEG contains a single DC-only scan. Because the DCT runs on 16x16 blocks, such an image won't a solid color: it'll be 1/16th of the original resolution. Scan number Channels DCT bins Precision 0 Y Cb Cr 0 - 0 Full Doing this, I can get Chrome to render around 90 frames before giving up. Other browsers like Firefox have more patience, but a 90 scan image seems to work almost everywhere. As a bonus, this avoids the ghosting of the naive attempt: that happened because AC scans are supposed to refine old data. Normally, this allows images to include multiple precision levels without inflating file size... but doesn't play nicely with my tricks. If the file only includes DC scans with no actual progression, this isn't a problem. Since a "DC-only" frame is a standards-compliant images , creating them doesn't require anything special: Using these, it's possible to pack a whole video inside a single image: Click to open in new tab Besides unconventional rickrolls and other trolling, this has no practical applications: there's no way to add timing information, so playback is entirely dependent on network delay. ... although there is a lot of fun to be had using partial rendering: This is a pure HTML video using <dialog> tags: badapple.rose.systems Of course, there's no rule that the data must be hardcoded: here's a interactive single-page application with no CSS or JavaScript. (seems slighty broken, I'll investigate later) Related : /projects/bad_jpeg/merge.c : The code used to generate these images /projects/bad_jpeg/merge.c : The code used to generate these images

0 views

Notes on the Fourier Transform

The Fourier series is a great tool for analyzing periodic functions. But what about functions that don’t repeat? We’ve seen that we can compute Fourier series for a non-periodic function defined on a finite interval, as long as we don’t care about its behavior beyond that interval. Let’s extend this idea to functions that never repeat; that is, non-periodic functions defined on the interval (-\infty,\infty) . To motivate the subject ahead, let’s look back at the example used in the earlier post about Fourier series : With an odd extension into [-2,0] . In that post, to make the Fourier series work, we assumed t(x) keeps repeating with a period 2L=4 on the entire x axis. Here, let’s face the reality that it does not - in fact - repeat, and observe how our Fourier series work out. Recall that the Fourier series approximating t(x) are the sine series (since it’s an odd function): The following visualization is interactive. By default, it shows t(x) (with its odd extension) and no Fourier series approximation. We’ll proceed by a series of steps and observe the outcome: Step 1 : set to some non-zero number; already at 3, the approximation is very good. The frequency spacing is \frac{\pi}{L} (this is the coefficient of x in the sines). Note that the Fourier series repeats every 2L , as expected. Step 2 : increase L to 6. This means our series are constructed assuming t(x) has a period of 12, not 4. Note how the Fourier series look now - they repeat every 12, and they don’t match t(x) as well as before. We can increase to a higher number to make the match better. As L grows, the spacing between adjacent frequencies decreases. Step 3 : increase L to 10. We no longer see the repetitions, so feel free to increase the values of x min and x max until you do. Note again that we need to add more and more coefficients to match t(x) better with this larger L , and the spacing adjacent frequencies grows smaller. Increasing L means our function repeats at larger and larger intervals. The logical conclusion of this progression is to ask - what happens if the function never repeats, meaning L\rightarrow\infty ? While not mathematically rigorous, the visual experiment here lets us make some conjectures: we’ll likely need an infinite number of coefficients for a good approximation, and moreover, the spacing between these coefficients will tend to zero. In other words, instead of a discrete set of coefficients, we’ll end up with a continuous line, or function . The function produced by this process is the Fourier transform of t(x) , and the next section shows its mathematical derivation. In these notes, we’ll be using the complex exponential formulation of Fourier series: We’re interested in a non-periodic defined on the interval (-\infty,\infty) . So we’ll be exploring the above equations for L\rightarrow\infty . First, let’s make a slight change of notation. Instead of writing formulae in terms of the period ( 2L ), we’ll be using the n-th harmonic angular frequency w_n : So we can slightly rewrite our series as: Using \Delta w as the difference between two consecutive frequencies: Using this notation, C_n is expressed as: So far there are no new insights here, just some new notation. Now we’re going to use it to facilitate the next step. Since L\rightarrow \infty , then \Delta w\rightarrow 0 . Let’s calculate the limit of the Fourier series representation of when \Delta w\rightarrow 0 : And substitute the latest C_n into this equation, changing its dummy integration variable from x to t to avoid confusion [1] Reordering slightly, and also replacing n\Delta w by w_n in the complex exponents: Looking at the limit with the sum carefully, this is a Riemann sum (see Appendix A)! w_n is the "sampled" version of , and \Delta w\rightarrow 0 . We can therefore replace it by an integral, changing w_n to and \Delta w to dw [2] : The inner integral is called the Fourier transform of and denoted [3] : And the full equation for is then the inverse Fourier transform: Let’s take our favorite odd triangular pulse example and calculate its Fourier transform. The function’s mathematical definition and plot are shown earlier in this post. Note that we’re not extending this function periodically - it’s zero beyond the range [-2,2] ; this is exactly why we need the Fourier transform here - as we’ve seen, Fourier series won’t do because the function they reconstruct eventually starts repeating. We’re looking to find: To calculate the integral, let’s decompose the complex exponent using Euler’s formula: Since our t(x) is odd, the first integral is zero . Also t(x)sin(wx) is even, so we can write: We’ve already calculated a very similar integral in the post on Fourier series , so let’s just skip to the result: The only remaining difficulty is its value at 0, which seems undefined at first (division by zero). However, note that as w\rightarrow 0 , the numerator also tends to 0, so we can use L’Hopital’s rule (twice!) to find that: This function is complex-valued; in fact, it’s purely imaginary. How do we visualize it? A common way to visualize complex-valued functions is by plotting their magnitude and phase separately. The magnitude of \hat{t}(w) is: Since \hat{t}(w) is purely imaginary, there are only two options for the phase: When the numerator is positive, we get a negative imaginary number with phase -\pi/2 , and when the numerator is negative, we get a positive imaginary number with phase \pi/2 . Finally, when \hat{t}(w)=0 (which happens at w=0 , by our earlier analysis, but also whenever is a whole multiple of \pi ), the phase is undefined. Here’s the magnitude and phase of \hat{t}(w) plotted against : It is common to talk about \hat{t}(w) as the frequency domain representation of t(x) . When the functions we’re working with have time as their domain (e.g. the x in t(x) represents time), which is often the case in the study of signals and systems, the Fourier transform can be seen as computing the frequency domain representation of the function. Here’s the Fourier transform formula again: It takes - the time domain representation of a function, and converts it to \hat{f}(w) - a frequency domain representation. For well-behaved functions, these two representations are dual - each one describes the function completely, just in a different way. To convert back from a frequency domain representation to the time domain, we use the inverse Fourier transform: While a time-domain plot ( t(x) ) shows how a signal changes over time, a frequency-domain plot ( \hat{t}(w) ) shows how the signal is distributed across all possible frequencies. Moreover, as we’ve seen, \hat{t}(w) is complex valued. Each frequency therefore has both a magnitude and a phase: the magnitude tells us how strongly that frequency contributes, while the phase tells us how that component is shifted. The frequency domain is extremely useful in signal analysis; for example, when designing filters. The Fourier transform also has a number of properties that are very useful in signal analysis and processing. But first, let’s discuss what a "well-behaved function" means for the purpose of applying Fourier transforms. The simplest existence condition for Fourier transforms is absolute integrability (also known as Lebesgue integrable): With this condition, \hat{f}(w) exists on the entire domain, is continuous and vanishes (tends to 0) as |w|\rightarrow\infty [4] . While this condition is sufficient, it’s not necessary; there are less well-behaved functions that also have Fourier transforms defined with some limitations. In these notes, we’re mostly interested in well-behaved functions that are used in real-world engineering, so we won’t discuss the other cases. Another assumption commonly made for real-world functions is that they vanish (tend to 0) as |x|\rightarrow\infty . While this is not a direct outcome of absolute integrability [5] , it’s a reasonable assumption in engineering. After all, real-world signals have finite energies. Intuitively, when we also assume is uniformly continuous , the assumption of vanishing at |x|\rightarrow\infty is a logical conclusion, because otherwise how can the total area for |f(x)| be finite? An important outcome of this discussion is that the Fourier transform is unsuitable for periodic functions. Functions that repeat at intervals are not absolute integrable . For periodic functions, we use Fourier series. The Fourier transform is a linear operator, because the integral is linear: So is the inverse Fourier transform; it’s similarly easy to show that: If we scale the domain of a function by a constant, its transform changes only slightly: Let’s do the variable substitution u=ax : This is the Fourier transform evaluated at \frac{w}{a} , so: There’s one small caveat here; when a is negative, the integral bounds should be flipped, causing a minus sign in front of the transform. So we can write: Which works for any a\ne 0 . This property is intuitive when thinking about signals: suppose a>0 , then f(ax) means the signal is compressed in the time domain by a factor a . The scaling property says that the frequency domain is expanded using the same factor; in other words, the higher frequencies become more prominent because we need sharper transitions to represent the compressed signal. Time shifting What happens to the Fourier transform if we time-shift the input signal by some constant: f(x-x_0) . By definition: Substituting u=x-x_0 , we get du=dx , so: Transform of a derivative An extremely useful property that’s often employed in the solution of partial differential equations; let’s calculate the Fourier transform of the derivative of : We’ll use integration by parts, where dv=f'(x) and u=e^{-i\cdot wx} . Therefore, v=f(x) and du=-iw\cdot e^{-i\cdot wx} : Recall the assumption made in the "Existence condition..." section about vanishing at infinities. So the first part of the equation above is zero, and we’re left with: Transform of convolution The convolution between two continuous functions and g(x) is defined as: Let’s calculate the Fourier transform of this function: This step of combining the integrals into a double integral, as well as the next step (changing the order of integration) is possible due to Fubini’s theorem and our assumption that and g(x) are Lebesgue integrable. Switch order of integration: Now, f(\xi) in the inner integral doesn’t depend on x , so we can pull it out: The inner integral is just the Fourier transform of a time-shifted g(x-\xi) , so we can write: And the remaining integral is the Fourier transform of , so: Convolution in the time domain translates to multiplication in the frequency domain! This result is so important in signal processing that it’s called the convolution theorem . Suppose we have some function and we want to know the area bounded between this function’s graph and the x axis in a certain interval [a,b] . One way to do this is to take a partition of the interval: And calculate the area under for every element of the partition. We can then approximate such sub-areas by rectangles, as follows: We’ll denote the area of each rectangle as f(x^*_i)\cdot\Delta x : There are many ways to choose which point of the interval [x_{i-1},x_i] to denote as x^*_i : left point ( x_{i-1} ), right point ( ), mid-point between the two (which is what our plot shows) or anything in between. The distinction doesn’t really matter for our purpose, as we will soon see. We can approximate the area under the curve of in the interval [a,b] with the Riemann sum , using a uniform partition: If is continuous on [a,b] , then as n\rightarrow \infty : This is known as the Riemann integral , or just the definite integral. The limit is why the exact choice of x^*_i doesn’t matter: as n\rightarrow\infty we have \Delta x\rightarrow 0 , and all points within [x_{i-1}, x_i] are equally good. The Fourier series is a great tool for analyzing periodic functions. But what about functions that don’t repeat? We’ve seen that we can compute Fourier series for a non-periodic function defined on a finite interval, as long as we don’t care about its behavior beyond that interval. Let’s extend this idea to functions that never repeat; that is, non-periodic functions defined on the interval (-\infty,\infty) . Visualizing Fourier series for non-repeating functions To motivate the subject ahead, let’s look back at the example used in the earlier post about Fourier series : \[t(x)= \begin{cases} x & 0 \leq x \leq 1 \\ 2-x & 1 < x \leq 2 \\ \end{cases}\] With an odd extension into [-2,0] . In that post, to make the Fourier series work, we assumed t(x) keeps repeating with a period 2L=4 on the entire x axis. Here, let’s face the reality that it does not - in fact - repeat, and observe how our Fourier series work out. Recall that the Fourier series approximating t(x) are the sine series (since it’s an odd function): \[t(x)=\frac{8}{\pi^2}\bigg[ sin\frac{\pi x}{2}-\frac{1}{3^2} sin\frac{3\pi x}{2}+\frac{1}{5^2}sin\frac{5\pi x}{2}-\cdots\bigg]\] The following visualization is interactive. By default, it shows t(x) (with its odd extension) and no Fourier series approximation. We’ll proceed by a series of steps and observe the outcome: n (terms in the Fourier series) L x min x max Step 1 : set to some non-zero number; already at 3, the approximation is very good. The frequency spacing is \frac{\pi}{L} (this is the coefficient of x in the sines). Note that the Fourier series repeats every 2L , as expected. Step 2 : increase L to 6. This means our series are constructed assuming t(x) has a period of 12, not 4. Note how the Fourier series look now - they repeat every 12, and they don’t match t(x) as well as before. We can increase to a higher number to make the match better. As L grows, the spacing between adjacent frequencies decreases. Step 3 : increase L to 10. We no longer see the repetitions, so feel free to increase the values of x min and x max until you do. Note again that we need to add more and more coefficients to match t(x) better with this larger L , and the spacing adjacent frequencies grows smaller. Increasing L means our function repeats at larger and larger intervals. The logical conclusion of this progression is to ask - what happens if the function never repeats, meaning L\rightarrow\infty ? While not mathematically rigorous, the visual experiment here lets us make some conjectures: we’ll likely need an infinite number of coefficients for a good approximation, and moreover, the spacing between these coefficients will tend to zero. In other words, instead of a discrete set of coefficients, we’ll end up with a continuous line, or function . The function produced by this process is the Fourier transform of t(x) , and the next section shows its mathematical derivation. Fourier series with L\rightarrow\infty leading to Fourier transform In these notes, we’ll be using the complex exponential formulation of Fourier series: \[f(x)=\sum_{n=-\infty}^{\infty}C_n\cdot e^{in\pi x/L}\] With: \[C_n=\frac{1}{2L}\int_{-L}^{L}f(x)e^{-in\pi x/L}dx\] We’re interested in a non-periodic defined on the interval (-\infty,\infty) . So we’ll be exploring the above equations for L\rightarrow\infty . First, let’s make a slight change of notation. Instead of writing formulae in terms of the period ( 2L ), we’ll be using the n-th harmonic angular frequency w_n : \[w_n=\frac{n\pi}{L}\] So we can slightly rewrite our series as: \[f(x)=\sum_{n=-\infty}^{\infty}C_n\cdot e^{i w_n x}=\sum_{n=-\infty}^{\infty}C_n\cdot e^{i\cdot n \Delta w x}\] Using \Delta w as the difference between two consecutive frequencies: \[\Delta w=w_n-w_{n-1}=\frac{n\pi}{L}-\frac{(n-1)\pi}{L}=\frac{\pi}{L}\] Using this notation, C_n is expressed as: \[C_n=\frac{\Delta w}{2\pi}\int_{-\pi/\Delta w}^{\pi/\Delta w}f(x)e^{-i\cdot n \Delta w x}dx\] So far there are no new insights here, just some new notation. Now we’re going to use it to facilitate the next step. Since L\rightarrow \infty , then \Delta w\rightarrow 0 . Let’s calculate the limit of the Fourier series representation of when \Delta w\rightarrow 0 : \[f(x)=\lim_{\Delta w\rightarrow 0}\sum_{n=-\infty}^{\infty}C_n\cdot e^{i\cdot n \Delta w x}\] And substitute the latest C_n into this equation, changing its dummy integration variable from x to t to avoid confusion [1] \[f(x)=\lim_{\Delta w\rightarrow 0}\sum_{n=-\infty}^{\infty}\left[\frac{\Delta w}{2\pi}\int_{-\pi/\Delta w}^{\pi/\Delta w}f(t)e^{-i\cdot n \Delta w t}dt\right]\cdot e^{i\cdot n \Delta w x}\] Reordering slightly, and also replacing n\Delta w by w_n in the complex exponents: \[f(x)=\frac{1}{2\pi}\lim_{\Delta w\rightarrow 0}\sum_{n=-\infty}^{\infty}\left[\int_{-\pi/\Delta w}^{\pi/\Delta w}f(t)e^{-i\cdot w_n t}dt\right]\cdot e^{i\cdot w_n x}\Delta w\] Looking at the limit with the sum carefully, this is a Riemann sum (see Appendix A)! w_n is the "sampled" version of , and \Delta w\rightarrow 0 . We can therefore replace it by an integral, changing w_n to and \Delta w to dw [2] : \[f(x)=\frac{1}{2\pi}\int_{-\infty}^{\infty}\left[\int_{-\infty}^{\infty}f(t)e^{-i\cdot wt}dt\right]\cdot e^{i\cdot w x}dw\] The inner integral is called the Fourier transform of and denoted [3] : \[\boxed{\hat{f}(w)=\mathcal{F}\left[f(x)\right]=\int_{-\infty}^{\infty}f(x)e^{-i\cdot wx}dx}\] And the full equation for is then the inverse Fourier transform: \[\boxed{f(x)=\mathcal{F}^{-1}\left[\hat{f}(w)\right]=\frac{1}{2\pi}\int_{-\infty}^{\infty}\hat{f}(w)e^{i\cdot w x}dw}\] Example calculation of Fourier transform Let’s take our favorite odd triangular pulse example and calculate its Fourier transform. The function’s mathematical definition and plot are shown earlier in this post. Note that we’re not extending this function periodically - it’s zero beyond the range [-2,2] ; this is exactly why we need the Fourier transform here - as we’ve seen, Fourier series won’t do because the function they reconstruct eventually starts repeating. We’re looking to find: \[\hat{t}(w)=\int_{-\infty}^{\infty}t(x)e^{-iwx}dx\] To calculate the integral, let’s decompose the complex exponent using Euler’s formula: \[\hat{t}(w)=\int_{-\infty}^{\infty}t(x)cos(wx)dx-i\int_{-\infty}^{\infty}t(x)sin(wx)dx\] Since our t(x) is odd, the first integral is zero . Also t(x)sin(wx) is even, so we can write: \[\hat{t}(w)=-2i\int_{0}^{\infty}t(x)sin(wx)dx\] We’ve already calculated a very similar integral in the post on Fourier series , so let’s just skip to the result: \[\hat{t}(w)=-2i\cdot\frac{2\cdot sin(w)-sin(2w)}{w^2}\] The only remaining difficulty is its value at 0, which seems undefined at first (division by zero). However, note that as w\rightarrow 0 , the numerator also tends to 0, so we can use L’Hopital’s rule (twice!) to find that: \[\lim_{w\rightarrow 0} \hat{t}(w)=0\] Therefore: \[\hat{t}(w)= \begin{cases} -2i\cdot\frac{2\cdot sin(w)-sin(2w)}{w^2} & w\neq 0 \\ 0 & w=0 \\ \end{cases}\] This function is complex-valued; in fact, it’s purely imaginary. How do we visualize it? A common way to visualize complex-valued functions is by plotting their magnitude and phase separately. The magnitude of \hat{t}(w) is: \[|\hat{t}(w)|=\sqrt{\hat{t}(w)\cdot\hat{t}(w)^*}=2\left|\frac{2\cdot sin(w)-sin(2w)}{w^2} \right|\] Since \hat{t}(w) is purely imaginary, there are only two options for the phase: When the numerator is positive, we get a negative imaginary number with phase -\pi/2 , and when the numerator is negative, we get a positive imaginary number with phase \pi/2 . Finally, when \hat{t}(w)=0 (which happens at w=0 , by our earlier analysis, but also whenever is a whole multiple of \pi ), the phase is undefined. Here’s the magnitude and phase of \hat{t}(w) plotted against : It is common to talk about \hat{t}(w) as the frequency domain representation of t(x) . The frequency domain representation of functions When the functions we’re working with have time as their domain (e.g. the x in t(x) represents time), which is often the case in the study of signals and systems, the Fourier transform can be seen as computing the frequency domain representation of the function. Here’s the Fourier transform formula again: \[\hat{f}(w)=\mathcal{F}\left[f(x)\right]=\int_{-\infty}^{\infty}f(x)e^{-i\cdot wx}dx\] It takes - the time domain representation of a function, and converts it to \hat{f}(w) - a frequency domain representation. For well-behaved functions, these two representations are dual - each one describes the function completely, just in a different way. To convert back from a frequency domain representation to the time domain, we use the inverse Fourier transform: \[\mathcal{F}^{-1}\left[\hat{f}(w)\right]=\frac{1}{2\pi}\int_{-\infty}^{\infty}\hat{f}(w)e^{i\cdot w x}dw\] While a time-domain plot ( t(x) ) shows how a signal changes over time, a frequency-domain plot ( \hat{t}(w) ) shows how the signal is distributed across all possible frequencies. Moreover, as we’ve seen, \hat{t}(w) is complex valued. Each frequency therefore has both a magnitude and a phase: the magnitude tells us how strongly that frequency contributes, while the phase tells us how that component is shifted. The frequency domain is extremely useful in signal analysis; for example, when designing filters. The Fourier transform also has a number of properties that are very useful in signal analysis and processing. But first, let’s discuss what a "well-behaved function" means for the purpose of applying Fourier transforms. Existence condition for the Fourier transform The simplest existence condition for Fourier transforms is absolute integrability (also known as Lebesgue integrable): \[\int_{-\infty}^{\infty}|f(x)|dx<\infty\] With this condition, \hat{f}(w) exists on the entire domain, is continuous and vanishes (tends to 0) as |w|\rightarrow\infty [4] . While this condition is sufficient, it’s not necessary; there are less well-behaved functions that also have Fourier transforms defined with some limitations. In these notes, we’re mostly interested in well-behaved functions that are used in real-world engineering, so we won’t discuss the other cases. Another assumption commonly made for real-world functions is that they vanish (tend to 0) as |x|\rightarrow\infty . While this is not a direct outcome of absolute integrability [5] , it’s a reasonable assumption in engineering. After all, real-world signals have finite energies. Intuitively, when we also assume is uniformly continuous , the assumption of vanishing at |x|\rightarrow\infty is a logical conclusion, because otherwise how can the total area for |f(x)| be finite? An important outcome of this discussion is that the Fourier transform is unsuitable for periodic functions. Functions that repeat at intervals are not absolute integrable . For periodic functions, we use Fourier series. Some useful properties of Fourier transforms Linearity The Fourier transform is a linear operator, because the integral is linear: \[\begin{aligned} \mathcal{F}\left[\alpha f(x)+\beta g(x)\right]&=\int_{-\infty}^{\infty}\alpha f(x)e^{-i\cdot wx}dx+\int_{-\infty}^{\infty}\beta g(x)e^{-i\cdot wx}dx\\ &=\alpha\int_{-\infty}^{\infty}f(x)e^{-i\cdot wx}dx+\beta\int_{-\infty}^{\infty}g(x)e^{-i\cdot wx}dx\\ &=\alpha\mathcal{F}\left[f(x)\right]+\beta\mathcal{F}\left[g(x)\right] \end{aligned}\] So is the inverse Fourier transform; it’s similarly easy to show that: \[\mathcal{F}^{-1}\left[\alpha\hat{f}(w)+\beta\hat{g}(w)\right]= \alpha\mathcal{F}^{-1}\left[\hat{f}(w)\right]+\beta\mathcal{F}^{-1}\left[\hat{g}(w)\right]\] Scaling If we scale the domain of a function by a constant, its transform changes only slightly: \[\mathcal{F}\left[f(ax)\right]=\int_{-\infty}^{\infty}f(ax)e^{-i\cdot wx}dx\] Let’s do the variable substitution u=ax : \[\mathcal{F}\left[f(ax)\right]=\frac{1}{a}\int_{-\infty}^{\infty}f(u)e^{-i\cdot \frac{wu}{a}}du\] This is the Fourier transform evaluated at \frac{w}{a} , so: \[\mathcal{F}\left[f(ax)\right]=\frac{1}{a}\hat{f}\left(\frac{w}{a}\right)\] There’s one small caveat here; when a is negative, the integral bounds should be flipped, causing a minus sign in front of the transform. So we can write: \[\mathcal{F}\left[f(ax)\right]=\frac{1}{|a|}\hat{f}\left(\frac{w}{a}\right)\] Which works for any a\ne 0 . This property is intuitive when thinking about signals: suppose a>0 , then f(ax) means the signal is compressed in the time domain by a factor a . The scaling property says that the frequency domain is expanded using the same factor; in other words, the higher frequencies become more prominent because we need sharper transitions to represent the compressed signal. Time shifting What happens to the Fourier transform if we time-shift the input signal by some constant: f(x-x_0) . By definition: \[\mathcal{F}\left[f(x-x_0)\right]=\int_{-\infty}^{\infty}f(x-x_0)e^{-i\cdot wx}dx\] Substituting u=x-x_0 , we get du=dx , so: \[\begin{aligned} \mathcal{F}\left[f(x-x_0)\right]&=\int_{-\infty}^{\infty}f(u)e^{-i\cdot w(u+x_0)}du\\ &=e^{-iwx_0}\int_{-\infty}^{\infty}f(u)e^{-i\cdot wu}du\\ &=e^{-iwx_0}\mathcal{F}\left[f(x)\right] \end{aligned}\] Transform of a derivative An extremely useful property that’s often employed in the solution of partial differential equations; let’s calculate the Fourier transform of the derivative of : \[\mathcal{F}\left[f'(x)\right]=\int_{-\infty}^{\infty}f'(x)e^{-i\cdot wx}dx\] We’ll use integration by parts, where dv=f'(x) and u=e^{-i\cdot wx} . Therefore, v=f(x) and du=-iw\cdot e^{-i\cdot wx} : \[\mathcal{F}\left[f'(x)\right]=\left[f(x)e^{-i\cdot wx}\right]^{\infty}_{-\infty}-\int_{-\infty}^{\infty}f(x)(-iw\cdot e^{-i\cdot wx})dx\] Recall the assumption made in the "Existence condition..." section about vanishing at infinities. So the first part of the equation above is zero, and we’re left with: \[\begin{aligned} \mathcal{F}\left[f'(x)\right]&=-\int_{-\infty}^{\infty}f(x)(-iw\cdot e^{-i\cdot wx})dx\\ &=iw\int_{-\infty}^{\infty}f(x)e^{-i\cdot wx}dx\\ &=iw\cdot\mathcal{F}\left[f(x)\right] \end{aligned}\] Transform of convolution The convolution between two continuous functions and g(x) is defined as: \[(f\ast g)(x)=\int_{-\infty}^{\infty}f(\xi)g(x-\xi)d\xi\] Let’s calculate the Fourier transform of this function: \[\begin{aligned} \mathcal{F}\left[(f\ast g)(x)\right]&=\int_{-\infty}^{\infty}e^{-i\cdot wx}\left[\int_{-\infty}^{\infty}f(\xi)g(x-\xi)d\xi\right]dx\\ &=\int_{-\infty}^{\infty}\int_{-\infty}^{\infty}e^{-i\cdot wx}f(\xi)g(x-\xi)d\xi\ dx \end{aligned}\] This step of combining the integrals into a double integral, as well as the next step (changing the order of integration) is possible due to Fubini’s theorem and our assumption that and g(x) are Lebesgue integrable. Switch order of integration: \[\mathcal{F}\left[(f\ast g)(x)\right]=\int_{-\infty}^{\infty}\int_{-\infty}^{\infty}e^{-i\cdot wx}f(\xi)g(x-\xi)dx\ d\xi\] Now, f(\xi) in the inner integral doesn’t depend on x , so we can pull it out: \[\mathcal{F}\left[(f\ast g)(x)\right]=\int_{-\infty}^{\infty}f(\xi)\int_{-\infty}^{\infty}e^{-i\cdot wx}g(x-\xi)dx\ d\xi\] The inner integral is just the Fourier transform of a time-shifted g(x-\xi) , so we can write: \[\mathcal{F}\left[(f\ast g)(x)\right]=\int_{-\infty}^{\infty}f(\xi)e^{-i\cdot w\xi}\mathcal{F}\left[g(x)\right]d\xi=\mathcal{F}\left[g(x)\right]\int_{-\infty}^{\infty}e^{-i\cdot w\xi}f(\xi)d\xi\] And the remaining integral is the Fourier transform of , so: \[\mathcal{F}\left[(f\ast g)(x)\right]=\mathcal{F}\left[f\right]\cdot\mathcal{F}\left[g\right]\] Convolution in the time domain translates to multiplication in the frequency domain! This result is so important in signal processing that it’s called the convolution theorem . Appendix A: Riemann sum and the definite integral Suppose we have some function and we want to know the area bounded between this function’s graph and the x axis in a certain interval [a,b] . One way to do this is to take a partition of the interval: \[a=x_0<x_1<\cdots<x_{n-1}<x_n=b\] And calculate the area under for every element of the partition. We can then approximate such sub-areas by rectangles, as follows: We’ll denote the area of each rectangle as f(x^*_i)\cdot\Delta x : \Delta x=(b-a)/n is the width of one interval (assuming a uniform partition, but the math works just as well for non-uniform ones). x^*_i is some value in the interval [x_{i-1},x_i] .

0 views
<antirez> 3 weeks ago

Control the ideas, not the code

Look at the past history of this blog. There are many blog posts about programming with AI, a few of them date back to January 2024 (like this: https://antirez.com/news/140). I’m a relatively well regarded programmer, after all. I don’t have the need to still be in the “loop” as a old man that seeks for relevance, I recently rejoined Redis, and now I also am developing a new open source software for local LLM inference that received a good welcome in the community. Why I keep doing this, of saying what people don’t want to hear? Why I keep announcing how future programming will be by default? Because I feel the urge of lowering the impact for people less prepared to the change than me, often younger than me, and that, unlikely me, didn’t see many of those things coming (In 2022 I published, before ChatGPT existed, a book preannouncing many things that now happened and other things that I believe *will* happen, so I feel like I can say this without sounding egocentric). So mine is a trick. People feel more and more programming is completely modified by AI and don’t know what they should do, if they can really start coding in a completely different way, without looking much at the code as their main output. They feel like they are betraying their own field. So my intention is to arrive and say “look at me, In can write code, you know, I’m not hiding behind AI: yet, things changed, it’s not your weakness, it’s not that you are AI-pilled. It is just that our field is evolving in an incredible *and* painful (but also joyful) direction”. This is why yesterday, on X, I said that I believe many programmers at this point have less impact they could have because they look at the code. I truly believe into that. And note that this does not mean to vibe code something just asking for the final product. The point is: if you control the ideas of your software, looking at the code itself is suboptimal and often pointless. For the following reasons: 1. You can now generate a lot of code, even *not* accounting for the LLM code verbosity (that is also effect of not being able to instruct them well, for most of the part). How are you supposed to review 5k lines of code every day? 2. LLMs are very good at writing locally optimal code, and are worse (but improving) with big ideas. What’s the point of scanning function by function, line by line? Instead you should prompt the design you have in mind, sometimes ask “how is exactly the design of that part? How does it work?”, and evaluate if it is the right model. It is much faster. 3. The working day is 8 hours. If you read the code, it is a tradeoff. You are doing less of what today is the most important part of your job, that is, asking yourself: what I’m doing with this software? What are the new directions I want to take? And also, think at new ideas, features, optimizations tricks. And doing a lot of QA. Controlling the ideas. Do you remember this phrasing from the Mythical Man Month? Well, a book from the 70s tells us more things about the current software era than many of the things that were said from 2000 to 2020. Why people that now protest against AI were not horrified by the state of software in the last decade? The level of slop we touched during recent years, before AI, is unbelievable. I’ll say you another thing. What is slop? With DwarfStar I implemented an inference for two LLMs (DeepSeek v4 and GLM 5.2) in a completely automated way, but: try it yourself, you will discover you can’t just say “implement XYZ” and see it working. You have to understand how things work, what is the best design, how to reach a certain level of performance. Then I compared the implementation, for correctness, to other systems, finding that other implementations sometimes contained more errors. I researched more, and found that the local inference world is full of subtle errors that accumulate and damage the model output, issues in the attention implementation causing performance slopes after the context is over a certain limit because indexed attention implementations are broken (do more work than they should, for instance), and so forth. This is the result of a domain that is very complicated to handle, fast changing, with models that are slightly different one from the other in the inference graph being released every day. It’s an unfair game for developers. Well: AI helps a lot with that. There are many domains where rigorous engineering (in the design side) and testing is *far* better than writing a GPU kernel by hand (or reading it). So are we sure most of that resistance it is not ideological? Matteo Collina yesterday asked me, in reply to my tweet: but didn’t you say that you check all the AI generated code for Redis? And this is a good question indeed. Yes, I do, but this is, at this point, something I *need* to do but that I believe to be mostly pointless, partially once GPT 5.5 was released, but now with Fable and GPT 5.6 Sol even more. Yes: I identify things that I don’t like how they are coded, but if I open other Redis files written by other Redis contributors there is *far worse*, and not since they are not good coders, but because it is a matter of taste. I write very clean code since I want it to be readable, so during the implementation of Redis Arrays I operated changes. I’m doing it again for the 50% memory saving optimization of Redis sorted sets, a PR that I’ll submit soon. But I do not feel this is useful anymore. Nobody should anymore look at this code, but only at the ideas the code contains. I continued to do it out of respect for users. Redis is at this point a commonly useful thing, and many programmers will open files and modify stuff by hand. But if I had my hands free, you know what I would do, instead? Use all the time that the review is taking me to do more QA, to think at the next optimization idea and apply it, and to use LLMs to write a DESIGN.md file where each data structure is described in human language, with the ideas it contains, the implementation tricks, the design. That, in the future, is going to be much more useful. Do you want to modify sorted sets? You open the file, read the design, then you own the ideas. You can open your agent and ask it what to do with the right mental model. This is a lot more useful than reviewing the code. Fable and GPT 5.6 reviews to the sorted sets memory saving are going to spot ways more errors and subtle race conditions that my review is going to uncover. Yet I’ll do it. But for the majority of software projects, all this does not make sense anymore. Focus on controlling the ideas, instead. Focus on quality, testing, and having an idea of the software you want to ship. The world changed and it is painful, but also full of opportunities to improve a software world that was already completely rotten. I have a doubt only regarding young programmers that don't have enough experience, and can't build a mental model. We don't know, yet, if they will require or not to understand very well how a given piece of code works, but I believe they should learn how to write programs. Yet, I'm not sure checking the LLM output is the right thing they should do. It may be a lot more useful if they learn some programming language and implement a small interpreter, a small database, an hash table and so forth. Reviewing some Javascript stuff of some web site for a customer? Hell, no, don't lose time with that shit. Comments

0 views
David Bushell 1 months ago

Astro is fine I guess

When I’m not fighting WordPress I deliver static HTML or the occasional JavaScript framework integration. For personal projects I have ‘fun’ with my own static site generator . This week was a side quest (soon to be main quest) to build my new company website. We’re talking proper business here so I can’t be messing about. I figured an off the shelf SSG would be most suitable. I asked the socials, “ 11ty or Astro ?” Both are popular but Astro had the edge. I gave Astro an early spin back in 2022 and found it slow . Maybe it’s good now? I ran with minimum release age to avoid immediately getting pwned . I selected Astro’s “Use minimal (empty) template” option and it generated both an and file — are you f — deep breaths, don’t fall for the rage bait. I code in a modern editor so I installed the recommended Astro extension. At first I struggled with Zed recognising HTML. I discovered a restart temporarily fixed the issue, but I guess I restarted one time too many because now the Astro LSP is completely broken. No modern comforts for me then. At least I can look at HTML without the red squigglies. I know what you’re going to say, “Dave bro, you’re inflicting this pain upon yourself! Just write HTML!” And I should. I just want native no-framework HTML includes , you know? Can you imagine the civilisation we’d live in if that could happen? I persevered and got my templates built with minimal fuss. I added a markdown collection and got the blog part blogging. It’s obvious that people use Astro to build real websites because all my “how do I” questions had an answer in the documentation. I’ve been forced to deploy way too many “React spaces” in my templates because Astro’s whitespace treatment is a mystery. I don’t need many components so I haven’t gone deep on Astro vs JSX . My site has zero JavaScript on the front-end. I plan to keep it that way. Edit: Christian Niklas on Mastodon shared a link to a recent Astro update where they added a option that defaults to no longer “following HTML rules.” Umm… okay. Set this to or if you’re building a website? I set it to . Minifying whitespace is over-optimisation. Astro has got the job done, despite the developer experience being broken out of the box. I dread to think what graveyard of dotfiles is installed if I choose a non-minimal start. I can easily de-Astro my templates should I need to. Right now Astro is solving the right problems and the issues are but a nuisance. Final conclusion: Astro is fine I guess. I’m not convinced Cloudflare’s acquisition is a good thing, considering their record for performative slop. I’ve lost my enthusiasm for DX and tooling to be honest. Even my own SSG experiments are collecting dust. I’d call the ecosystem a lost cause if I was being dramatic. I just try to avoid the worst of it and care about the end product: shipping a damn fine website! Which I can’t do because I’ve got more businessing to business before this particular site sets sail. Maybe in a few months? It’s looking awesome on though. Thanks for reading! Follow me on Mastodon and Bluesky . Subscribe to my Blog and Notes or Combined feeds.

0 views
Ankur Sethi 1 months ago

Data locality (sometimes) beats algorithmic complexity

I've been ECS -curious ever since I learned about it in the Bevy game engine documentation . The ECS architecture predictably improves performance in languages that give you low-level control over memory (C, C++, Rust, Zig, and friends). But how does it fare when used in high-level, dynamic, garbage-collected languages such as JavaScript? This is the question Dan Murphy set out to answer in The Physics of Memory : Is it possible to use an ECS-style architecture in Javascript? And for applicable operations, does that actually do better than objects + V8’s garbage collection? To answer the question, Murphy built a 2D physics simulation of 15,000 balls bouncing around in a box using several different techniques. He found that a JavaScript implementation of the simulation that used ECS outperformed the usual "giant graph of objects" OOP implementation by 24x. He writes: It's also worth noting how the usual OOP implementation creates GC pressure: In OOP, entities are scattered across the heap. As they move and interact, the JavaScript engine’s garbage collector is constantly triggered, and the CPU frequently stalls waiting for pointer lookups. This causes sporadic frame drops (micro-stutter). Because ECS uses pre-allocated, flat TypedArrays, memory access is 100% predictable and GC overhead is zero, guaranteeing perfectly smooth frame delivery. My favorite thing about Murphy's post is that you can run all his benchmarks in your own browser. I love it when technical explanations or benchmarks are accompanied by embedded "apps" you can play around with. I'm surprised at how much data locality matters for performance. An algorithm with worse big-O complexity can outperform one with better complexity if it makes good use of the CPU's L1/L2 caches. Very cool. Cache Locality > Algorithmic Complexity : At 15,000 entities, pointer-chasing and unpredictable tree branching cannot compete with the contiguous L1/L2 cache locality of a flat 1D array sort—even though trees have a better theoretical Big-O complexity. You Don’t Need WASM for ECS Wins : Simply switching your JavaScript codebase to a flat Structure of Arrays (SoA) layout yields up to a  24x speedup  over OOP. WASM is the cherry on top (another 2.5x), not the entry ticket. Pragmatism Wins : While a hand-tuned SoA is the absolute fastest, using a production ECS library like   still gives you a massive  14x speedup  over OOP while providing a clean, scalable API. IMO, for 99% of applications using a library is the correct engineering choice.

0 views
Simon Willison 1 months ago

The new GPT-5.6 family: Luna, Terra, Sol

OpenAI's latest flagship model hit general availability this morning , and comes in three sizes: Luna, Terra, and Sol (from smallest to largest). The new models are priced per 1M input/output tokens as Luna $1/$6, Terra $2.50/$15, Sol $5/$30. For comparison, the Claude Opus series are $5/$25 and the Claude Fable 5 is $10/$50, but price-per-million tokens doesn't tell us much now that the number of reasoning tokens can differ so much between models for the same task. All three models have a February 16th 2026 knowledge cutoff, a million token context window, and 128,000 maximum output tokens. OpenAI's biggest benchmark claim concerns long-running agentic performance, with one benchmark showing all three models outperforming Claude Fable 5: We trained GPT-5.6 to get more useful work from every token. On Agents’ Last Exam , an evaluation of long-running professional workflows across 55 fields, GPT-5.6 Sol sets a new high of 53.6, eclipsing Claude Fable 5 (adaptive reasoning) by 13.1 points. Even at medium reasoning, it beats Fable 5 by 11.4 points at roughly one-quarter the estimated cost. That efficiency extends to smaller models, which are essential to making intelligence more abundant and affordable: GPT-5.6 Terra and GPT-5.6 Luna outperform Fable 5 at around one-sixteenth the cost. Amusingly, one self-reported benchmark that Fable 5 crushed the GPT-5.6 family on was SWE-Bench Pro, where Fable 5 got 80% compared to GPT-5.6 Sol getting 64.6%. This may help explain why OpenAI chose to publish this article yesterday specifically calling out SWE-Bench Pro for problems they found while auditing that benchmark: In light of these results, we estimate that ~30% of SWE-bench Pro tasks are broken, and advise that model developers carefully examine results I've had some early access to GPT-5.6 Sol - it's definitely very competent, though so far it hasn't struck me as better than Fable at the kind of complex coding tasks I've been using with Anthropic's model. As usual, the model guidance for using GPT-5.6 has the most interesting details. There are a bunch of new API features that I need to explore (and probably add support for in LLM ), including: Here's a full page with 18 different pelicans - for reasoning efforts none, low, medium, high, xhigh, and max across the three different models. It also lists their token and calculated costs - the least expensive was gpt-5.6-luna at effort none for 0.71 cents, the most expensive was gpt-5.6-sol at max reasoning level for 48.55 cents. In further pelican news, if you jump to 17:50 in their livestream from this morning you'll see OpenAI's own demo of 3D pelicans riding a tricycle, a bicycle, a pony, and another pelican! You are only seeing the long-form articles from my blog. Subscribe to /atom/everything/ to get all of my posts, or take a look at my other subscription options . Programmatic Tool Calling allows the models to "compose and run JavaScript that orchestrates tool calls" - which sounds to me like it could help bridge the gap between MCPs and full terminal sessions that can compose CLI utilities in useful ways. Also reminiscent of the dynamic filtering mechanism Anthropic added to their web search tool, which allows code execution against web results as part of a single model turn. Multi-agent lets the model "spin up subagents for parallel, focused work" - the sub-agent pattern now baked into the core API. Prompt cache breakpoints brings the Claude model of prompt caching to OpenAI, letting you be explicit about where the cache breakpoints are rather than relying on the API to detect them automatically. Personally I much prefer automatic detection (still supported by OpenAI), but presumably there are optimization cost savings to be had here if you put the work in. You can now set detail: original on image requests to avoid resizing the image at all before it is processed.

0 views
Farid Zakaria 1 months ago

Who does Anubis actually stop?

I have been working on a patch to the Linux kernel to support for the interpreter ( ) via bpf in [ thread ]. Of course I’m leveraging an LLM to help me do this! To pre-seed the context of the LLM, I asked it to read the https://lore.kernel.org/ thread. Uh oh. Looks like they have adopted Anubis , which is an HTTP proxy that requires proof-of-work before allowing access to the resource. Did this really do anything? Unfortunately, no. My AI diligently came up with anubis-fetch , which you can find at https://github.com/fzakaria/anubis-fetch . The tool tries to natively solve the proof of work or, as a last resort, will launch Chromium to visit the URL. This tool also impersonates a real Chrome TLS/JA3 fingerprint natively via req so it clears passive Cloudflare blocking too. ☝️ So who did we stop? The exact adversary Anubis targets defeats it trivially. The whole use of Anubis feels regressive and marginalizes those without access to “good” AI. For a scraper, solving the Anubis challenge is a one-time, amortized-to-zero cost since the cookie can be cached and reused. For a human, it’s seconds of spinner, battery drain on every fresh visit. They can’t amortize anything amongst each other. This “regressive tax” is paid even more so by those with weaker devices or who access the content on their phone. Clients that don’t leverage JavaScript (e.g., text browsers (w3m/lynx), screen readers, RSS readers) are completely left out. Did deploying Anubis stop any of the aforementioned bot-farms or are they mildly inconvenienced when they had to augment their bots to support a new proof of work solution briefly? The irony is that Anubis’s goal is to stop AI but it was incredibly easy for AI to circumvent it and yet the cost to humans and an open web remains. With the presumption Anubis is now a regressive tax, how much does it cost us? Every number here is a rough estimate. This is not a environmental argument at all since the bot-farmers and AI tools themselves are using many orders of magnitude more energy. Nevertheless, it’s interesting to see how much time is spent doing proof-of-work challenges that marginalize people. Difficulty is the number of leading zero hex characters the hash must have, so the expected work per solve is hashes. Difficulty 4 is the common default. Rates assumed: ~50 MH/s native (Go), ~0.5 MH/s in-browser JS; “felt” wall-clock includes page load, the worker, and the reload. Let be the number of Anubis challenge-solves per day, worldwide. Assume a felt time of and device energy per solve (screen + CPU). Collectively we are wasting an impressive amount of time waiting for access to websites; time we didn’t spend before the AI era. As a human, time is precious and finite to me, whereas to a robot it is not. Human-time / year = Energy / year (kWh) =

0 views
Takuya Matsuyama 1 months ago

Inkdrop Roadmap vol.6: Completed 🎉 — Now preparing for the official v6 release

Hi folks, it's Takuya here, the solo developer of Inkdrop . I'd like to report a status update on the Inkdrop project here. About a year and a half ago, I published the roadmap of Inkdrop vol.6 . And I'm happy to announce that every planned feature and improvement on that roadmap is now done! 🥳 They all shipped as part of the v6 canary series — 21 canary releases so far, built and tested together with the community. When I wrote the roadmap, I honestly wasn't sure how long it would take. I would have been surprised if the me of that time had seen this result. Thank you so much for all your feedback along the way — I couldn't have done it without you. Even beyond the roadmap, I've added so many new features and improvements. So, I'm confident you'll enjoy it if you're coming from v5. Let's dive into what I accomplished along the roadmap, what came out of it beyond the plan, and what's next. What made the development slow down was the huge technical debt, as I mentioned in the past post . Inkdrop was originally built on the Atom editor's framework, and when Atom was sunsetted in 2022, many of the modules it depended on were no longer maintained. I had to replace them one by one while keeping the app stable — the hardest and least visible part of this journey. With v6, that debt is finally paid off. Here's a quick before & after: None of these are shiny features on their own. But they're exactly what allowed me to ship everything you'll see below, and they make Inkdrop much faster to develop going forward. The codebase is now modern, healthy — and honestly, fun to work on again. I'm an indie developer, and Inkdrop is a one-person project — so manpower has always been the bottleneck. Paying off the tech debt was a particularly big headache: some of the inherited modules were so large that it originally took the whole Atom team to maintain them. But thanks to the recent advancements in coding agents, that burden finally feels manageable — and even enjoyable to tackle. AI didn't just speed up the coding; it changed how I work: These new workflows have opened up possibilities that simply didn't exist for solo developers before. A refactoring of this scale used to be unthinkable for one person — now I can maintain a codebase that once took a team, and spend the saved energy on what matters most: the product itself and my users. Here's the roadmap vol.6, item by item, with what actually shipped: The roadmap was only half the story. While working through it, I ended up rebuilding a huge part of the app and shipping a lot of features that weren't planned. Here are the highlights, grouped by area: And on top of all that, hundreds of bug fixes reported by canary testers. The community has also been building amazing plugins on the new APIs — note-tabs (browser-like note tabs), code-runner (run JS/Python code blocks in notes), constellation (an interactive note graph), copy-as-jira , kanso-ink (theme), and more. Existing plugins are getting v6 support too, like hitahint , link-compact , thumbnail-list , and editor-utils . My goal remains the same as I wrote in the roadmap: keep improving the core user experience without bloating the app, so you can stay focused on taking notes. I believe v6 embodies exactly that. You can download the binary here: Please create a topic on the “ Issues > Canary ” category. This is the most preferred way for me because I can manage which issue has been resolved or not. We have our Discord server , where you can casually discuss and talk with other users. With the roadmap completed, I've shifted gears to preparing for the official release of v6 . That means polishing the details, stabilizing the canary builds, updating the documentation and the website, and helping plugin and theme authors migrate. Especially, building a new landing page is gonna be fun! I'm also going to work on the mobile app as well. The official v6 release is getting close. Stay tuned! 💪 I manage implementation plans as Inkdrop notes and let the agents work through them. Watch: Note-driven agentic coding workflow using Claude Code and Inkdrop I built and published a tool to manage multiple Claude Code sessions on tmux . While building the AI features, I had an agent explore Zed's source code and save the report to Inkdrop , to learn how it implements similar functionality. ✅ Share target & share extension — You can quickly stock web pages into Inkdrop from other apps on mobile. ( v5.5.0 ) ✅ Command palette — It became Telescope , a versatile Spotlight-like search bar (the name is borrowed from telescope.nvim, haha). It fuzzy-searches commands, notebooks, tags, and the table of contents of the current note, with scope prefixes like for commands and for notebooks. It's extensible, so plugins can add custom sources. ( canary.1 ) ✅ Migrate to CodeMirror 6 — The biggest one. The whole editor was rebuilt on CodeMirror 6, and it enabled a bunch of new editing features: a floating toolbar, slash commands, GitHub Alerts syntax support, emoji autocompletion, autocompletion inside code blocks, and quick note-link insertion with . ( canary.1 ) ✅ Outline view — Powered by Telescope. Click the button in the editor header (or run ) to jump between sections. It highlights the current section based on your cursor or scroll position, and even lists task items. It's provided as a plugin ( telescope-toc ), which doubles as a reference implementation for custom Telescope sources. (Thanks Basyura-san for the original sidetoc plugin!) ( canary.6 ) ✅ Preview pane improvements — Copy buttons for code blocks landed in both the preview and the editor, and double-clicking an image opens it in an image viewer. As a bonus, find-in-preview finally works — it highlights matches even across DOM elements, which is essential for finding text in code blocks. (Thanks q1701 and Basyura for the original plugins!) ( canary.2 , canary.4 ) ✅ Two-factor authentication — OTP-based 2FA is available for your account. ( v5.11.0 ) ✅ Prepare for ARM64 & other platforms — This required repaying a lot of technical debt. I replaced the deprecated LevelDB backing store with SQLite , stopped bundling (which used to bundle all of Node.js and npm!), and rebuilt it as a lightweight standalone CLI ( @inkdropapp/ipm-cli ). As a result, Inkdrop now supports ARM64 on Windows and Linux , plus Flatpak and AppImage packages for modern Linux distros. ( canary.1 , canary.4 , canary.5 ) ✅ Improve image upload speed — Attachments are now uploaded in parallel via signed URLs, so syncing image-heavy notes is significantly faster. ( canary.12 ) ✅ Diff view for revision history on desktop — The diff view I loved on mobile is now on desktop, too. ✅ Notebook icons — You can assign custom icons to notebooks from a picker with 1,500+ icons from the Lucide icon set, with category tabs and search. Icons show up everywhere — the sidebar, Telescope, and notebook selectors. ( canary.9 ) ✅ Visualize your progress and achievements — The activity stats view shows how many notes you created and tasks you worked on over the past 52 weeks, along with your current and longest streaks. Note-taking is a contribution to your work, after all! ( canary.14 ) ✅ AI integrations — Shipped as an opt-in, bring-your-own-API-key design, so you stay in control of your data. The inline AI assistant transforms selected text in place with built-in prompt presets (proofread, summarize, Mermaid diagrams, Markdown tables, and your own custom prompts). Next Edit Suggestions predicts your next edit like GitHub Copilot — set to manual trigger by default so it doesn't distract you — and it can even draw context from your linked notes and backlinks. ( canary.16 , canary.18 , canary.20 ) Reading highlights — Select text and hit the highlight button to wrap it in a tag, rendered beautifully in the preview. Perfect for emphasizing what resonates in your reading notes. ( canary.3 ) Native spellcheck support — The editor now uses the OS-native spellchecker. ( canary.10 ) Smarter link pasting — Pasting a URL now suggests link formats inline through the autocompletion menu instead of a dialog, and the page title is fetched in the background so nothing interrupts your flow. ( canary.15 ) Create a note from autocomplete — Start typing a title after , choose "Create new note," and it's created, linked, and opened in one step. ( canary.16 ) Little things that add up — ToDo item strikethrough, link-open tooltips, commands (Thanks Lukas and TheRabidOstrich !), View menu toggles for line numbers / line wrapping / readable line length, and a refurbished editor header with navigation back/forward, view mode buttons, and a native action menu (Cmd/Ctrl+J). ( canary.2 , canary.3 , canary.12 , canary.18 ) Embed GitHub code snippets by pasting a link — Paste a GitHub source URL and the code is fetched and inserted as a syntax-highlighted snippet with line numbers and a link back to the source. Connect your GitHub account via OAuth and it works with private repos too, including rich link titles for repos, issues, and PRs. ( canary.6 , canary.11 ) Advanced code blocks — Language icons, line numbers, and meta info rendering, plus GFM highlighting inside fenced code blocks — nested code blocks and YAML frontmatter included. ( canary.6 , canary.9 , canary.20 ) Mermaid got a serious upgrade — A pan & zoom toolbar with a full-screen viewer, and diagrams are now themed entirely through CSS variables, so they automatically match your theme in light and dark mode. (Thanks @inkwadra for the original pan/zoom PR!) ( canary.21 ) Manual notebook ordering — Drag and drop notebooks in the sidebar into your preferred order; it syncs across devices. ( canary.9 ) Fuzzy matching everywhere — Telescope, the notebook and tag list menus, and the tag input all use the same fuzzy-matching algorithm, so you find things fast without spelling them right. ( canary.15 ) Quicker navigation — Filter buttons for notebooks and tags in the sidebar, a search bar in the notebook picker, context menus on the workspace and note-list headers, and a sort-order button that shows the current order as a label. ( canary.6 , canary.15 , canary.16 ) Keep running in the system tray (Windows & Linux) — Handy if you use the local HTTP API, and it makes reopening the app instant. (Thanks Kyoichiro-san and Micha for the request!) ( canary.21 ) Plus a custom-built tooltip UI, a macOS "Look Up Selection" context menu, and an account usage stats tab. ( canary.14 , canary.16 ) A new CSS-variable-based theming system — Themes are now a thin layer of variables over the base styles instead of a full Semantic UI stylesheet, which makes them far easier to build and maintain. ( canary.18 ) One theme package instead of three — The UI / syntax / preview theme types inherited from Atom have been merged into a single unified package that styles the whole app. ( canary.21 ) Live theme previews — The Themes preferences show preview cards rendered live from each theme's color palette, and is uploaded to the plugin registry to power previews before you install. ( canary.20 , canary.21 ) New official themes — Kanagawa ( Wave / Dragon / Lotus ), Solarized ( Light / Dark ), and Nord ( Dark / Light ), plus a default syntax theme overhaul built on modern CSS like . ( canary.18 , canary.20 , canary.21 ) Dropped Electron's module — I replaced it with type-safe IPC bridges in a massive architectural overhaul. Database access from plugins became roughly 13x faster , and the app is more secure because only intended methods are exposed. ( canary.11 ) SQLite as the backing store — Replacing the long-deprecated LevelDB unblocked ARM64 support and repaid one of the oldest debts from the Atom era. ( canary.4 ) Modern build pipeline — Migrated from Webpack + Grunt to electron-vite (Vite + Rolldown), which made production builds 10x faster and the dev build launch almost instant. I also converted all Less stylesheets to plain CSS, moved drag & drop from the unmaintained to , and kept Electron riding the latest releases throughout the canary series. ( canary.14 , canary.18 ) Security hardening — Access keys moved to the system keyring, and the login flow is protected with Cloudflare Turnstile against credential-stuffing bots. ( canary.16 , Security Update ) A brand-new CLI — No more bundled Node.js and npm. It publishes tarballs directly like npm (no more committing compiled files to GitHub), and scaffolds a new plugin or theme in seconds with TypeScript all wired up. ( canary.5 , canary.18 ) Official TypeScript definitions — @inkdropapp/types gives plugin authors full type safety without exposing the app's internals. ( canary.14 ) Auto-installed essential plugins — mermaid, math, and markdown-emoji are installed and kept up to date automatically, and you can disable them anytime. ( canary.14 ) Vim plugin improvements — Relative line numbers (Thanks @p1n9_d3v !) and an option to keep Vim registers separate from the system clipboard (Thanks @birtles !). ( canary.11 ) Updated docs — The plugin migration guide and theme development guide are refreshed for v6, along with new component and module references. https://my.inkdrop.app/download/canary Inkdrop Website: https://www.inkdrop.app/ Send feedback: https://forum.inkdrop.app/ Join the Discord server: https://docs.inkdrop.app/start-guide/join-discord-server 𝕏: https://x.com/inkdrop_app 🦋: https://bsky.app/profile/devaslife.bsky.social

0 views