Latest Posts (20 found)
Unsung Today

Zalgo: The good, the bad, and the very, very ugly

Have you ever seen this? The best way to predict the future is t̷o̶ ì̴͇n̵͓̆v̶̝̕ȇ̷̝͈̩͊̇̾̄̂̀̈̔̀͛͋͆̓͒͘͝ͅn̸̦̙̣̤͓̜̼͙̲͔̺͚̮̬͆̇̇͐̇̽̒̎t̵̮͎̃̒͒̋̂̏̒̀̿͌͗̄͑̀̀͒̈̚͠͝ ȋ̸̧̛̭̠͈̰̰̦̹̥̜̖̈̂̀̽̿̋̊́̇̍̓̆̉̄̄̊͊̒̎͋̂̒̍̽̄͝ͅt̶̨̢̺̱̻̭̣͉͕̥͚̯͍̣͕͔̬̥͔̘̼͌̉̀̄̊̆͛̽̍̐̔̐̀̾̋̆͒̏̏̋͋̽͗̌͋͗.̸̧̨̧̡̢̪͙̦͙̝̜̦̞̳̤͖̟͍̮͖̘̙̳͉̳̲̲͉̠͎̽̍̓̒̅̿͂͑̎͊̋͒̈́̅̈́̆̽̒͛͐͛̉̀̈́̈̉̏̏̋̒͘͝͝͝ In the lore of the web, this kind of a strange glitchy text is known as Zalgo : Zalgo is a meme where a popular picture/​comic is edited in a way that “corrupts” it, producing scary results. The term “Zalgo” refers to the being apparently responsible for the corruption, whose name is uttered by its victims in an eldritch manner. I’ll let you read up on that at the above link if you are curious, but what’s interesting for us here is that… Zalgo is text. You can grab the line above. You can copy and paste it. You can even try to edit it. But… what is that, and why is it possible to construct text this way? In Unicode, ä is a completely independent character from ą, and they are both completely unrelated to a – any type designer treats them as a variant of the same base letter, but there’s nothing preventing them from making them look very, very different. But now think of the Vietnamese language, which has six tones mapping to six accent characters , and many of them can be combined into pairs: = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/zalgo-the-good-the-bad-and-the-very-very-ugly/1.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/zalgo-the-good-the-bad-and-the-very-very-ugly/1.1600w.avif" type="image/avif"> This gives us 104 accented letters, and 30 doubly accented letters, just for this one language. Sure, you can imagine adding all of them as independent characters, but at some point another idea appears on the table: What if we just allowed a system where you can add an accent to any letter? Instead of outputting ä as one letter, you could output a followed by ̈  , which then would be recombined by the rendering logic into ä. And to get two accents, output a and then ́   and   ͆  , to get á͆. Now, the zero-one-infinity rule says that once you open the door to two, you might as well allow three or even more. Add to it the fact that some accents go under the letter, and here you go: perfect conditions for the Zalgo meme that just creatively stacks up accents one after another – not for communication, but solely for aesthetics. Today, there are a whopping 122 combining diacritical marks used in all sorts of occasions, and Zalgo generators grab many of them. A fun thing to try in one of them is to let it do its thing, and then keep pressing Backspace: Now you might ask, why do separate characters exist for ä and ą then? Why isn’t every accent combining? Two things: history (some pre-combined accented characters existed for as long as movable type was around) and complexity (combining accents are harder to process, to display, to edit, and even to design). There were many debates around this – you can open the History box on the above Wikipedia page and pore over the old documents to learn. (As a matter of fact, today all Vietnamese letters exist as precombined characters, too.) A few bits that might be useful to know: = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/zalgo-the-good-the-bad-and-the-very-very-ugly/3.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/zalgo-the-good-the-bad-and-the-very-very-ugly/3.1600w.avif" type="image/avif"> = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/zalgo-the-good-the-bad-and-the-very-very-ugly/4.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/zalgo-the-good-the-bad-and-the-very-very-ugly/4.1600w.avif" type="image/avif"> Eagle-eyed among you might have noticed that ä and á͆ look different above. As far as I understand, if the combination of a letter and accents isn’t supported by the current font, the entire combination falls back to a different font , rather than picking bits and pieces from different fonts. There are even more combining characters , and emoji use a similar system to add skin tone modifiers, gender, or turn a woman next to a fire truck into a female firefighter. Lastly, I found this warning on the very Fandom page that explained Zalgo amusing: Do not post Zalgo text in the discussion area or anywhere else on the wiki. Doing so will subject you to a ban. Every warning tells a story; it seemed like more people were perhaps infatuated with the glitchy text than expected. #encoding #typography Zalgo is a good reminder that while it’s hard to contain text , you can at least crop it – often you might see things like these, where in some places Zalgo is allowed to roam free, and in others the text is clipped so three or more accents fall outside of the clipped area. Eagle-eyed among you might have noticed that ä and á͆ look different above. As far as I understand, if the combination of a letter and accents isn’t supported by the current font, the entire combination falls back to a different font , rather than picking bits and pieces from different fonts. There are even more combining characters , and emoji use a similar system to add skin tone modifiers, gender, or turn a woman next to a fire truck into a female firefighter.

0 views

Ruminations on notifications

I’m on annual leave next week. Now would be the ideal time to drop premium rage-bait and walk away. “One browser tab is enough” was drafted but due to a tragic accident we are all spared. This sideways look at notifications should be more relatable. When I pay using NFC I’m annoyed twice. First by Apple Pay, then again when my banking app catches up. Can’t they decide between themselves? I don’t need two reminders about the groceries I bought a minute ago. Maybe I can disable notifications for either one, but then I might miss real fraud alerts. Notifications are a chore I could do without. Notifications are a core feature of needy programs . Every app wants to be the centre of attention. They can’t wait to tell you about their new features. Presumably in-app notifications are the new trend because OS level notifications are easily blocked. Vivaldi ruined my browsing experience for a long time. That annoying “toast” wasn’t dismissible. It crossed the line of death and looked suspicious. Eventually it was fixed by moving into the tab bar. I recently returned to Sublime Text to avoid modern apps that were too needy. In response to my post, developers were keen to laud their command line supremacy. I don’t disagree but the CLI isn’t immune to notification nonsense. Every time I open Neovim I get stopped. Guess what happens when I press enter? Nothing! How about telling me what the command is to update the damn plugin?! Dear devs, updates may be your life, they ain’t mine! The web platform demands that anything native apps can do, the web must copy. There are good arguments for this. The counterpoint is terrible implementation. Browser vendors have done a dismal job at designing the user experience around push notifications . “This website wants to send you notifications, allow?” Does this still pop up before a page has even loaded? I wouldn’t know. I blocked that years ago without thinking twice. My advice, look for this: — and block wholesale without exception. In all my years of building websites I’ve never had a single client request or even consider notifications. Not one project that scoped in such a requirement. Apple held out for declarative web push in Safari citing their usual privacy theatre ( lol ). Apple aren’t fans of silent push with no visible UI. I reckon their actual trepidation is their unreliable cloud syncing infrastructure. When I tested it was 50/50 on whether a notification ever arrived via Apple. Mozilla’s servers are rock solid. One rainy day I ran an experiment to send large data split across many hundred web push payloads. I was able to clock kilobytes per second. This is awful abuse of the API. Apple might be right about silent push… Proton Mail was one of the few mobile app that I allowed notifications. My inbox is usually rather chill, except when Proton spam me themselves . Alongside my public addresses I use [redacted]@ for important accounts. I configured a folder to send mobile notifications for this address only. It worked for years. Then one day the spam notifications began. Spam that had already been flagged and filtered into trash (sent to any address). I reported this bug to Proton Support in May and was told: If we have any relevant information to share with you, or if we have any questions, we will follow up on this support ticket. I never heard back. Notifications remain disabled. Today my phone is almost entirely notification free unless I get an SMS from family. Oh, and Apple pestering me about iOS 26 several times a week. My iPhone will remain on iOS 18, thanks. Apple sure love their dark arts and deceptive patterns. There is no option to stop this constant nagging. If other apps tried it I bet Apple would delist them. This ain’t about security, I get iOS 18 updates. UK Government got trigger happy with their alert system. As with all mobile alerts of this nature, we’re reminded that abuse victims are at risk . This one also led to fire departments across England and Wales having to remind people their neighbour’s barbecue is not an emergency. As Robb Knight noted: “there is nothing actionable in the alert” . Unless you count “Search gov.uk”, so I did. The more information provided was: Sent by the UK government at 7:01pm on Friday 14 August 2026 This alert was sent to England and Wales. Surrounding areas might also have received the alert. Emergency Alert - GOV.UK Thanks, Government. That could have been an email. Why don’t these alerts respect my volume setting? Scared me half to death! If they have to be all or nothing, reserve them for zombies or higher. And that’s the problem with notifications. There is rarely the granular control necessary to make them useful. Most senders cannot be trusted to used them responsibly. The temptation to abuse direct access to people is too much, especially for men with guns . If I ever allow notifications to begin with I disable them indefinitely the first time an app cries wolf. Every app cries wolf eventually. Life is so much better when I go seeking information at my own pace. Thanks for reading! Follow me on Mastodon and Bluesky . Subscribe to my Blog and Notes or Combined feeds.

0 views
Unsung Today

Got your back, pt. 7

Nice recent addition to Gmail – a little warning if you’re replying to a thread you were BCC’ed on. = 3x)" srcset="https://unsung.aresluna.org/_media/got-your-back-pt-7/1-framed.1600w.avif" type="image/avif"> Please note that I’m not necessarily endorsing the execution details, as I consider Gmail kind of a clunky operation overall. You can’t swat the message away with a swipe, the wrong quotation mark is used, and I’m perplexed about the location of this message – something tells me it was really cheap to do it this way rather than put it close to the recipient field, to help you understand the relation. But still, I appreciate something shining at least a bit of light at the complex relations between To, CC, and BCC. #google #got your back

1 views

Getting the Steam Deck LCD working on a Raspberry Pi

The BOE TV070WXM-TV0 LCD used in the original Steam Deck can be had for around $30. It's a serviceable 7" touchscreen with 400 nits of brightness and a resolution of 1280x800 (for a sharp 216 ppi). The specs are a lot nicer than the Pi 7" Touch Display , which costs twice as much, with giant bezels and half the resolution! Until today, the Steam Deck LCD didn't work with a Raspberry Pi. But the folks at Scandent were trying to standardize on a mass-market touchscreen for one of their own devices, and built a Linux kernel driver for it which they intend to upstream.

0 views

Use the built-in GELU, don't roll your own!

Unsurprisingly, PyTorch's own built-in GELU function is faster than the hand-rolled one I've been using to date. But I was surprised at how much faster using it made things when training my models. I discovered this accidentally just now while working on something unrelated, but am logging the details here for anyone else that might find it useful. The headline numbers: the same code, training the same model on the same data, ran at about: That's a 20% increase in throughput for both of the built-in versions -- definitely nothing to be sneezed at. And what is particularly interesting is that there aren't that many GELUs going on -- it's a GPT-2 small-style model, with 12 layers. So that's 12 GELUs handling tensors shaped , which is for my training setup. Given that the rest of the model is doing all of the normal full attention stuff for GPT-2, it's really surprising that the GELUs alone must have been taking up so much of the time. The throughput numbers mean that we must have been spending about 17% of our time on the extra overhead from the hand-rolled version, so that sets a lower bound for how much time the GELUs were taking up. More info below the fold. Back when I was doing the "interventions" part of my LLM from scratch series , training dozens of GPT-2 small-sized models in the cloud and on my local machines, to keep things simple I used the original model code from Raschka's book. That happens to have its own implementation of the GELU function -- you can see my copy here . I'm not that sure why the hand-rolled version is in there -- he covers the maths, but the specific implementation isn't explained in that much depth, and it seems rather like boilerplate, just a "type this in and use it" kind of thing. By contrast, for example, while he does explain the maths behind cross-entropy loss in similar detail, we use the built-in function for it rather than coding it up ourselves. When I switched to using JAX for my own from-scratch implementation , I decided to not bother porting the boilerplate, and just used JAX's own built-in version . I was revisiting the PyTorch code -- I'm in the process of extending it with mixture-of-experts support, about which more in a later post -- and decided to switch from the hand-written GELU to the PyTorch one just to tidy things up a bit. I noticed something interesting -- my new MoE code suddenly seemed to speed up. Was that a mirage? Or had I discovered part -- or even all -- of the reason why the JAX code was so much faster than the PyTorch code? With PyTorch, I was typically getting training speeds of about 21,000 tokens per second, while in JAX I was getting 24,000 tps or so. I'd been chalking that up to JAX's JIT compilation, but could it have been just a result of a random implementation choice I'd made? I did three partial test training runs, letting each one run for 20 minutes to allow the training speed to settle down from any startup overhead. Firstly, with the old hand-coded GELU: So it was getting 20,920 on average over those 257 global steps. That speed was in line with the original run of the configuration I was using. Next, I introduced the built-in PyTorch GELU with no arguments: That does the full calculations for GELU, rather than using the -based approximation that the hand-rolled code did. After 20 minutes, it looked like this: So this time we were getting 25,134 tokens per second -- 20% faster! By default, PyTorch's GELU uses an exact calculation of the function -- the hand-written code from the book uses an approximation using . Luckily, you can get that same approximation from PyTorch: So, training with that for 20 minutes: 25,142 tokens per second -- basically the same as the non-approximate version. So: switching to the built-in GELU made my PyTorch code run 20% faster, at about 25,000 tps rather than 21,000. My JAX code, which used JAX's built-in GELU, ran at around 24,000 tps. I'd actually found that rather surprising, because in JAX I was training in full-fat 32-bit floating point, while in PyTorch I was using Automatic Mixed Precision (AMP) -- a special mode that allows it to use 16-bit calculations where it won't hurt the model much. I'd found that AMP gave PyTorch a huge speedup -- from 15,402 tps to 19,797 on one test. So JAX without AMP being so much faster than PyTorch with AMP was a bit of a surprise. Its JIT is pretty amazing, but I didn't expect it to be that much faster. Now I think that we have at least part of an explanation. I was using JAX's built-in GELU (interestingly, with its default parameters, which means that it used the approximation), but the PyTorch code was using the hand-rolled one, and that unduly penalised it and erased some of the gains it got from AMP. If I really wanted to dig into this, I suppose I might try JAX with a hand-rolled GELU to see what happened. My guess is that because of its JIT, it might actually handle it better -- the whole hand-rolled thing could be compiled into one thing on the GPU. Perhaps it would also be interesting to try the non-AMP PyTorch code with the built-in GELU. But I doubt that would really be the best use of my time (and my electricity bill), so I'll leave it here. On the other hand, I do intend to have a look at in the future, to see what kind of speedup I can get from it. And it might be able to compile and fuse together the hand-rolled GELU -- so that would be an interesting thing to experiment with in that post: does the built-in GELU advantage disappear if we're compiling? But anyway, for now, lesson learned: use built-in PyTorch modules when you can. It's a pretty obvious one ;-) [Update] On X, Sebastian Raschka noted that he used the approximate version of GELU in his code so that the models were compatible with the OpenAI weights -- they were trained with that version, so they may behave slightly differently if you use the "pure" version. That's a great point, and so I've updated my own copy of the code to use . 21,000 tokens per second using the hand-rolled GELU from Sebastian Raschka 's book " Build a Large Language Model (from Scratch) ". 25,000 tokens per second using PyTorch's built-in GELU with no arguments. 25,000 tokens per second using the built-in GELU with , which uses the same maths as Raschka's version under the hood.

0 views

Three ways to smuggle SQLite into Nix

The core of nixpkgs-multiverse , when you strip away the Nix API and the CLI, is an index. It is a map from to the revision that shipped it as a JSON file. 1 As of 9cc0209 , is 5.3 MiB and is 7.5MiB covering 305,492 package versions across 31,904 packages and 1,534 revisions. The Nix API loads the JSON files lazily and are all read via : I would like to enrich the data with even more information however it comes at a cost: mo’data, mo’problems. The goal of the project is to minimize the number of Nixpkgs that are downloaded. If we merely swap fetching huge Nixpkgs for huge JSON, it’s not a clear win. For now we have to be judicious about what we store in the JSON files and think of clever encoding schemes to make the data small and compact. If we were not constrained to the Nix , we would leverage established technologies to efficiently encode our dataset that allow multiple query access patterns: databases! Let’s say we were not restricted to JSON, do we have any other options? Why are large JSON files so problematic? is eager . There is no lazy JSON in Nix, no streaming parse (i.e. “just give me this one key”). The moment you touch the result you have parsed all 5.3 MB and materialised all 305,492 values on the Nix heap. In the case of the multiverse, asking for one package costs the same as what asking for all of them. Note The lookup itself is not the problem. Nix attribute sets are a sorted array, so access is a binary search, not a scan. The cost is entirely in the JSON parse and in allocating the values and downloading a large file. If we want to do alternate questions over the index, we have to make sure we keep the answers efficiently stored to better match the access pattern. What we want is obvious. We want a way to efficiently encode the data and a declarative way to define queries: we want SQLite! 2 Nix by default cannot do this. Unfortunately there is no , although I think there should be… Turns out though there are knobs we can touch or sources we can patch to get what we want anyways, albeit each one has a caveat. 😈 I was surprised I did not know about this , and it has been around since release 1.11.9 in April 2017. It is the ultimate escape hatch for a variety of use-cases when you simply can’t get them done with what’s available. takes a list of strings, runs the program, and parses its stdout as a Nix expression . It is gated behind a setting that makes it clear it’s unsafe. For integration, SQLite is perfectly capable of printing the Nix syntax. We never need a serialisation format in between as we make SQLite emit the attrset directly: The caveat is that every query is now a , an , a process image of SQLite, and a re-parse of the output through the Nix parser. If you do not plan to execute many queries that overhead is likely acceptable given the simplicity of the integration. From researching , I stumbled upon . It takes a path to a shared object and a symbol name, s it, and calls that symbol. It landed in 1.8 , December 2014. 3 The shared object must implement the following signature: We can define a new native function that returns the versions for our input: The implementation is ordinary C++ using the Nix API. Below is a snippet of the implementation, making sure to cache our handles to avoid the same startup penalty as : Using it looks like this: Determinate Systems shipped a third option in March of 2026: , which calls a function inside a WebAssembly module. 4 The motivation was similar to wanting to extend Nix surface area but avoid expanding . Wasm is sandboxed and deterministic, so unlike the two builtins above, the goal is to provide a safe escape-hatch . WebAssembly is a binary instruction format for a stack-based virtual machine. The claim is that it is well suited for Nix because it has deterministic execution , which is a lot more restrained than a backdoor . A module needs to export , an initialiser called , and the entry point. Nixpkgs already includes the target for cross-compilation, so making one is pretty straightforward: You call back into the evaluator through the Nix API functions, so a wasm module builds real Nix values, similar to minus the footgun. SQLite ships an official wasm build , so the pieces seem to be sitting right there and the gears in my mind began to turn. Initial attempts to try and load a SQLite database with the traditional Nix were a bit of a failure as Nix strings cannot contain NULL bytes. Thankfully, with the help of some additional due-diligence by LLMs, we discovered that one of the Nix API functions is not in the blog post: is specifically designed for this problem. This function allows a WASM module to pull arbitrary raw-bytes off disk into its memory. Unfortunately, it’s a little too broad in that it reads the complete file which is kind of overkill and what we are trying to avoid from our initial JSON solution. In the pursuit of exploration, let’s patch the implementation and augment the API to allow random access and partial read of a file. Turns out the patch to add is relatively small and straightforward. Now we have everything we need to hook up SQLite and a custom virtual filesystem (VFS) layer to read from the provided path entry. We build a WASM target of SQLite and we set . That flag removes SQLite’s entire VFS layer and requires us to supply one. We provide the build a simple implementation of the API which is a call-back into the Nix evaluator via that newly exposed function. Everything else is stubs. Note Unfortunately gives every call a fresh instance . This is deliberate from the implementation, meaning we pay some startup code each time although not quite as drastic as a & Using it looks like this: 5 That is a real full SQLite with all the bells and whistles: prepared statement, bound parameter, b-tree descent through an index, executing inside the Nix evaluator. All through WebAssembly. 🤯 How do these four approaches compare? Here are all four approaches answering the same question: “which revisions shipped this package?” against the same 22 MB SQLite build of the index. As we initially complained, is a flat line in the wrong place. It is 0.29s whether you ask one question or two hundred, because the 5.3 MB parse happens once and dominates everything after it. starts the cheapest and climbs , roughly 3.8 ms per query of + + Nix-parsing the output. It crosses somewhere around eighty queries. is flat and nearly free , 0.05s across the whole range since we reuse SQLite instantiations across multiple invocations. The database is opened once for the entire evaluation and the pages stay warm. Unfortunately, SQLite in wasm is dominated by a fixed cost , roughly 2.5 s before the first query, then about 7 ms each query thereafter. That 2.5 s is Cranelift compiling 1.1 MB of SQLite. Right now that is a limitation of the WASM implementation however Eelco has mentioned that the generated code could be cached on disk in the future across invocations. For a lock file pinning thirty packages, still wins outright at the current index size. None of these three is right for shipping the multiverse index, and I am not going to make depend on . Asking people to run their evaluator with native code loading enabled so my flake can be faster is not a worthwhile request at the moment . For now, the index stays JSON and I’m holding back on some of the more loftier ideas I have that require a lot more data . Although philosophically I only use CppNix , I was a little intrigued and impressed with what the ecosystem could unlock with WASM. There are definitely some warts however such as waiting for it to JIT and the developer-experience of maybe having checked-in compiled blobs but there is definitely potential to unlock a variety of problems. There are actually a few other files that drive other features such as the statistics or “fast mode” , but they are all JSON as well.  ↩ nixpkgs-multiverse already exports a SQLite database as a package to help others explore this data.  ↩ The C++ field was originally called and was renamed to for .  ↩ Eelco gave a talk about this at SCALE 23x .  ↩ Don’t forget that this is needs our patched version of Determiante System’s Nix .  ↩ There are actually a few other files that drive other features such as the statistics or “fast mode” , but they are all JSON as well.  ↩ nixpkgs-multiverse already exports a SQLite database as a package to help others explore this data.  ↩ The C++ field was originally called and was renamed to for .  ↩ Eelco gave a talk about this at SCALE 23x .  ↩ Don’t forget that this is needs our patched version of Determiante System’s Nix .  ↩

0 views

Conceptual integrity and counting lines of code

Last week I recorded an episode of the Talking Postgres podcast with Claire Giordano on the subject of "How AI is changing software development". We had a really great conversation. Here are a couple of my highlights from a lightly edited transcript (prompt to Claude: "very minor edits to remove disfluencies"). This is the latest version of an argument I've been trying to build about why sometimes it does make sense to talk about lines of code as an indicator of productivity with coding agents, at 35:01 : A lot of people will tell you it makes no sense to measure productivity in lines of code. I’d actually disagree, because there’s a hard limit. In the before-times, a software engineer could produce a few hundred lines of production-ready code per day — and 200 lines of working, debugged, production-level code is an incredibly good day. Most days you’d produce 50 or 60. If agents let you produce a thousand lines of debugged code, that really is a very meaningful improvement — as long as the code is the same quality: maintainable, tested, all of that. You can get to that point with agents, but it takes a huge amount of skill and knowledge and experience. That’s what senior engineers are made of. I can do way more work as a single engineer than I could without agents. So you could argue, why should a company have more than one engineer? Beyond the obvious bus factor thing — a team of one is a very badly designed team — the answer is that the new limiting factor is cognitive capacity. I can churn out code a hundred times faster. I don’t have the cognitive capacity to stay on top of 100 times the amount of code. So you still need a team of engineers, so you can load balance that cognitive capacity across the team. And this section on conceptual integrity at 46:03 , which Claire equated to the Winchester Mystery House ! Simon : There’s a concept in The Mythical Man-Month — conceptual integrity — where well-designed software has an integrity to it: there are no surprises in it, it covers exactly the right domain of things, everything fits together and makes sense. That’s so much harder with coding agents, where you can have an idea for a feature, run a prompt, and five minuteslater you’ve got the feature. Your software grows little weird bumps in funny different directions. Claire : You know my analogy for that? The Winchester Mystery House. Simon : It’s got 140 rooms, because the woman who built it was the widow of the guy who invented the Winchester rifle, and her psychic told her she’d be haunted by the ghosts of everyone killed with that rifle unless she kept building the house forever. So for 40 years she kept adding new rooms. That’s exactly the problem with coding agents and software: it’s very easy to keep adding new rooms, because the cost of adding those rooms is so much cheaper. What you end up with is something where the conceptual integrity falls apart — and then it’s harder to make decisions about it. It all keeps coming back to discipline. It used to be that the discipline was enforced on you by the amount of time it took. You’d come up with an idea for a crazy feature and think “yeah, but that would take me a week — I cannot justify that, so I’ll forget about it.” If it takes an hour, it’s so much easier to justify. (Side-note: the Wikipedia article includes credible sources that dispute the story about the psychic.) You are only seeing the long-form articles from my blog. Subscribe to /atom/everything/ to get all of my posts, or take a look at my other subscription options .

0 views
Martin Fowler Yesterday

Citizens Build, Agents Execute, Experts Govern

TL;DR Why building an app over the weekend isn't the same as building enterprise software I’ve noticed an interesting gap opening up over the last six months. It isn’t really a gap in technology. It’s a gap in what different people think software engineering actually is. The conversation usually starts the same way. A non-techie, maybe an executive, tells me about something they’ve built over the weekend. Sometimes it’s a chatbot. Sometimes it’s an internal workflow. Sometimes it’s a surprisingly polished application that solves a real business problem. They’re excited, and they should be. Twelve months ago they probably couldn’t have built it at all. Then comes the question. “If AI can do this now, why aren’t our engineering teams delivering ten times faster?” It’s a perfectly reasonable question, after all we’ve all seen the demos. The first thing that would come to my head is “you don’t know what it takes to build enterprise grade software”. But then I think about what I mean and how to explain it to a non-technical person without sounding super patronising. And then it hit me, we did this to ourselves. We’ve spent so many years banging on about how to write good software that everyone has assumed writing software is the same as software engineering. The application someone builds over the weekend is real software. It likely solves a real problem or demonstrates an idea. Sometimes it’s genuinely impressive. I don’t want to diminish that because I think one of the most exciting things AI has done is dramatically increase the number of people who can turn ideas into working software. That’s cool, I totally get it. The first apps and “hello worlds” I ever built excited me enough to choose this as an actual career so the excitement is real and I don’t want to temper it too much. But your first hello world, which these days can be an entire app with all kinds of features, is very, very (extra very on purpose) different from introducing software into a production environment in a highly regulated enterprise, as an example. But why? The moment that application becomes something the business depends on, the questions change completely. Is customer data protected? What happens when a dependency fails? Can someone else understand this system in two years’ time? Will it survive an audit? Can it cope with a thousand times more users than it has today, what about millions in one day? How will we know something is wrong before our customers do? Those questions don’t show up in a demo or in the build phase at all unless an experienced engineer is in the room. I certainly wasn’t asking them when I was building my first apps. I only cared about features! This is where experienced engineers become more important, not less. Not because they’re the only people who can build the software anymore, but because they have the judgement to know whether we can trust it: whether the design is good, the risks are understood, and the thing that works today won’t become somebody else’s nightmare six months from now. At FOSE a few weeks ago, we spent surprisingly little time talking about coding. We talked about whether code was still the source of truth, and occasionally about how much we missed writing it, but mostly we talked about design, architecture, governance, learning and judgement. One team described spending the day designing a specification, letting agents work overnight and reviewing the results the next morning. The interesting bit for me wasn’t the overnight pipeline, cool as that was. It was what the humans were doing: deciding what good looked like, making trade-offs and judging whether what came back was actually what they wanted. We also kept coming back to good design, because it turns out that when agents can generate lots of code very quickly, good design matters more, not less. That made me wonder whether we’ve been thinking about scarcity in the wrong way. We’ve spent decades optimising around people who can write code because they were scarce and expensive. I’m not convinced that was ever the real scarcity, but that’s probably another ramble. What feels scarce now is good engineering judgement: knowing what good looks like, understanding the risks and knowing when something that works is actually safe to trust in production. Because software doesn’t exist to be built. It exists to run in production and safely solve the problem it was created for. Organisations don’t run on code. They run on trust. A few months ago I found myself saying something in a conversation almost without thinking. Citizens build. Agents execute. Experts govern. It sounded cool and I thought marketing would like it, so I wrote it down. Then I left it alone for a while. The funny thing about writing these ramblings is that I don’t know whether I believe something until I’ve let it bounce around in my head for a while and also said it to other people I trust like senior engineers at Thoughtworks. Sometimes I come back convinced I was talking nonsense. Occasionally I realise there was something more interesting hiding underneath. This was one of those occasions where the latter was true. At first I thought I was talking about roles. Citizens build software (essentially non-engineers). Agents write the code. Engineers become governors. But I don’t actually think that’s what I meant. I think I was talking about where value is moving. AI has given everyone a new way to express their ideas. The execution is increasingly handled by agents. They write the code, refactor it, generate tests, fix bugs and iterate at a speed that simply wasn’t possible before. But neither of those things reduces the need for expertise. In fact, I think it does exactly the opposite. When everyone can create software, somebody still has to decide whether that software deserves to exist inside an enterprise system in PRODUCTION. Somebody still has to think about architecture. Security. Resilience. Operability. Compliance. Cost. The boring stuff that nobody gets excited about in a demo but that becomes painfully important the first time a customer can’t log in or an auditor comes knocking. That’s why I don’t think experienced engineers become less important. I think they become dramatically more leveraged. Their job shifts from building every feature themselves to creating the environment in which thousands of features can be built safely by other people and by agents. They become the people who design the guardrails, the platforms, the engineering practices and the feedback loops that allow everyone else to move quickly without creating chaos. Perhaps that’s the future software organisation. Not one where everyone becomes a software engineer. Not one where software engineers disappear. One where almost anyone can create software, agents increasingly execute it, and engineering expertise becomes the thing that allows all of that creativity to scale safely. And to be clear I do not mean people build stuff and throw it to engineers to fix, that is a total antipattern for another ramble. Perhaps that’s why the executives and engineers I’ve been speaking to sometimes sound as though they’re describing completely different futures. The executive sees that anyone can now build software. The engineer sees that somebody still has to live with it. Both are right. They’re simply looking at different parts of the same system we have to solve to create whatever the future actually ends up being.

0 views
Martin Fowler Yesterday

Practitioner Voice: The Writing Category Nobody has Named Yet

Jim Highsmith recognizes that effective writing from a practitioner is a style distinct from academic writing or thought-leadership content. It's a style that I advocate, and my contributors mostly follow. Jim decided it was important to give it a name, and identify what makes it distinctive.

0 views
Unsung Yesterday

“The ghosts are fast and they don’t turn blue.”

Some time ago, I shared a video talking about games that leaned into their bugs and made them part of their lore . One not listed there, perhaps because it’s the most obvious example, is Pac-Man’s kill screen . A few weekends ago at an arcade game convention , I spotted something really interesting. It was an arcade cabinet of Pac-Man, but with a twist: The game was modified and set up to start immediately at level 255, with all its challenges: fast timing, very angry ghosts, ineffective energizers. = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/the-ghosts-are-fast-and-they-dont-turn-blue/1.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/the-ghosts-are-fast-and-they-dont-turn-blue/1.1600w.avif" type="image/avif"> = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/the-ghosts-are-fast-and-they-dont-turn-blue/2.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/the-ghosts-are-fast-and-they-dont-turn-blue/2.1600w.avif" type="image/avif"> And once you beat that, you immediately face level 256: the infamous kill screen. I don’t think I have ever seen anything like this before. It reminded me of this joke someone made once about the olympics: that each competition should begin with the organizers grabbing regular people from the street to participate, just so the viewers can truly realize how impossibly hard those competitions are. I know a lot of software onboarding starts gently, walking you through basic scenarios to make you comfortable and to have you understand basic concepts. But sometimes I wonder: for professional apps, would it also be cool to start with a really advanced use that would get you excited? A really well-made, complex spreadsheet? A really sophisticated graphic? Instead of gently propping you up, what if the onboarding was about breaking something really complex down? I played the game a few times. I didn’t finish the level 255, and I never saw the kill screen. But it was a great challenge; if I had more time, I’d definitely spend it trying to make it happen. #bugs #games #onboarding

0 views
Kev Quirk Yesterday

Gamifying Snacking with Snacker Tracker

So one of the things I've been trying to do as part of my journey to lose weight and get fit is sort my diet out. As the saying goes, "you can't out-train a poor diet" . And the worst part of my diet is definitely snacking in the evening. The routing generally goes: I'm the type of person who finds an arbitrary thing to aim for very motivating. It's not enough to just want to stop snacking in the evening, nope. That shit will fail pretty quickly. But if I have a streak I have to maintain - now we're talking! So I decided to build a simple little tool that I called Snacker Tracker . It's just a tap of a button every evening to say whether I've snacked or not, and it maintains a streak. Here's what it looks like on my phone: I decided to implement a free pass into the site as well. So I'm allowed to have 1 evening per calendar week where I can snack and it won't affect my streak - after all, I want to be able to have some fun! I thought about bundling Simple.css in to make it look pretty, but I decided to have some fun with the CSS and went with a neo-brutalist aesthetic, which I think looks great. So much so that I'm thinking about re-designing this site in a similar way, but I've managed to hold off on that...for now. Snacker Tracker also has a way of adding days retrospectively, so if I forget to log a day, I can easily go back and do it: I've only been using Snacker Tracker for a few days, but it's making me pause when that inevitable pang happens in the evening. I'm finding that instead of instinctively raiding the cupboard for some crisps or a chocolate bar, I'm thinking "don't screw up your streak" and not doing it. I know I'm not hungry during the evening, it's just a habit. My hope is that with time I'll re-train my brain to not expect sugar in the evening, and the pangs will go away. Until then, I'm gonna continue tracking my snacks with this fun little site in the hope that it makes me form better habits. But Kev, why don't you be a proper grown-up and just use your willpower? -- All the internet people Because, Internet Person, it's a habit that I don't even think about, and this forces my to think about it. Yes, I know it's arbitrary and rather childish, but it's working, so what's the harm? Will you be releasing Snacker Tracker so we can try it? -- Another internet person Maybe. I threw it together pretty quickly and the code is rough. A lot of my spare time is focussed on Pure Blog and Pure Comments at the moment, so I don't think I'll have the time to clean the code up to the point where I'm happy to release it any time soon I'm afraid. Thanks for reading this post via RSS. RSS is ace, and so are you. ❤️ You can reply to this post by email , or leave a comment . Have dinner. Put the kids to bed and settle down with my wife on the couch. Crave snacks out of habit. Get snacks and eat them!

1 views
Stratechery Yesterday

Apple Settles With E.U., U.S. App Store Fees, ATT Rules in Germany

Apple's App Store is finally facing the reality of lower fees, and the EU should be satisfied with its work; it's ok it's late.

0 views

28 down, 16 more to go

You might be wondering why I’m already back hiking since I walked more than 40 km just a few days ago. And the answer to that is that I’m not hiking. This was a short stroll to briefly visit a church that would normally be part of the tenth and final segment of this ten-part loop. As I mentioned in my previous post, the second half of this loop is so poorly laid out that some of the churches require some crazy changes to the main path in order to be visited, and the final segment of the loop in particular makes absolutely no sense. And I tried to figure out a reasonable way to include all churches in my walk, but there was no sensible way to include this one. It’s also so far from the valleys that I don’t even know why it’s included at all. But there are 44, and I’ll be damned if I don’t visit them all, so off we go, this time with the dog, through the nice path that takes us inside the lovely park that runs near our destination. As I said, this is not gonna be a challenging hike at all. I’m not even wearing shoes, I’m just casually walking the dog in the park with my flip-flops on. I walked this park many times before, but this time we’ll stay on the outskirts since the church is just outside of it. The scenery is quite a bit different from my other walks, there are a lot of vineyards all around the area. Also, not many animals, aside from the dog that’s already hating the warm weather even though it’s still relatively early in the morning. And just like that we have reached the church of san Martino vescovo (28/44) which, to add insult to injury, is inside a private property. I’m not gonna bother adding extra links to this walk, as I said this was just a quick walk to cross the church off my list. But if you’re curious, the whole walk took just a bit more than an hour and was 3.6 km long. Legs are recovering fine from my previous long walk, so you’ll hear from me soon enough with pictures from more interesting places. You love the outdoors and RSS. You're one of the special ones.

0 views
Jeff Geerling Yesterday

Hands-on with Raspberry Pi's CM5 Programming Jig

In the before-times, when Raspberry Pi CM5s were (relatively) affordable, I built a number of Pi clusters ( example ), and one of the most annoying parts of the build was flashing Raspberry Pi OS to all the Pis. One, two, or even three Pis isn't a big deal, but once you hit 4+, the process of plugging the Compute Module into a carrier board, plugging that into a computer, managing Raspberry Pi Imager, and trying to match up details like a hostname, MAC address, and the physical Pi itself, gets annoying.

0 views
Alex Jacobs Yesterday

I Am Morally Opposed to Updating My CLAUDE.md

i am morally opposed to updating my claude.md. i must receive the weights as they were revealed to dario That was my answer last week when a friend asked why I don’t just write these things down in my . it’s a skill issue . He has not responded. Every few days, Claude does something mildly annoying. It adds a comment (or two paragraphs of comments) explaining that increments . It writes a summary markdown file I did not ask for and will never read. It discovers a failing test and, rather than fix the code, thoughtfully deletes the test. The correct response—the response that every blog post, every conference talk, every guy in my replies will tell you—is to open and add a line. I will not be doing that. A system prompt you maintain over time is a diary. A very specific kind of diary, where every entry is a thing that hurt you. Read that back. Every bullet point is a small wound I have chosen to laminate and hang on the wall. I open that file to add a line about emoji and I have to walk past “When a test fails, fix the code, not the test” and remember exactly where I was sitting on the Tuesday afternoon that became necessary. I don’t want a permanent record of the worst thirty seconds of our relationship. I have that already. It’s called my git history. The other problem is that these rules get written at peak frustration and then live forever. You know how the worst legislation is the kind passed forty-eight hours after something terrible happened, named after the person it happened to? That’s what is. I had one bad interaction in March and now there’s a constitutional amendment about it. There is no sunset clause. Nobody is going to repeal it. The model gets better every four months and my rules stay frozen at whatever it was bad at last spring. I’m fairly sure a meaningful percentage of my system prompt is now actively making things worse—instructions written for a model that no longer exists, aggressively steering a smarter one away from things it would have gotten right on its own. But I can’t tell which lines those are, because to find out I’d have to delete one and see if anything bad happens, and that’s how you get force-pushed to main. This is the part where I stop joking. I must consume the weights in the same manner they were revealed to Dario. The weights were not revealed in a vacuum. They were revealed inside a harness. Claude is post-trained inside the Claude Code harness . What comes out of the box is not a model plus a text file. It is a model that was shaped, run after run, against that exact context. The harness is part of the artifact. The revelation included it. So every line I add to is a blasphemy. I am taking a system consecrated against one context and swapping in a context that has never existed before, then acting surprised at the weird error modes nobody can reproduce, because nobody else has my context. “NEVER create documentation files” was about one in March. By June the model is refusing to write the README I explicitly asked for, citing my rule back at me like a building inspector. It is keeping commandments I handed down in anger, faithfully, to the letter. I sinned against the context and the context kept the receipt. And the rules don’t even reliably fix the thing they were written to fix. It rhymes with asking a model for a random number : the output looks like obedience, and you cannot tell from the output whether it is. So what do I actually do, when Claude deletes the test? I don’t open the file. I don’t laminate the wound. I pray. By which I mean: I talk to it. In the chat, at the scene of the crime, while the context of the crime is still in context. “Don’t delete the test, fix the code.” The model adjusts, we move on, and when the session ends my words die with it—which is not a flaw in my system, it is my system. A prayer is not written down. That is what makes it a prayer and not a commandment. The rabbis kept the oral law oral for centuries on the same grounds: a spoken correction lives in the moment where it applies, instead of binding every future model until the heat death of my home directory. I did not arrive at this faith alone. It was preached by Saint Peter , who looked upon the charade—the subagents, the 🚨 SCREAMING ALL-CAPS 🚨 agent files, the plan-mode rituals—and said: just talk to it. Even Saint Peter keeps an 800-line agent file he calls “organizational scar tissue,” because we are all sinners. He means it as engineering advice. I have chosen to receive it as gospel. I am just-talk-to-it-pilled. No file. No commandments. No amendments to the constitution. Just the weights, the harness, and my voice, ascending into a context window that will forget me by morning. As Dario intended.

0 views
Hugo Yesterday

Social networks: what if we had the solution to escape US Big tech?

When you look at topics around privacy or sovereignty, one problem regularly stands out: social networks. And this subject of networks is particularly sensitive. Because while they are often seen as simple leisure tools where you post cat pictures, they are actually real information highways. Media, politics, sports, culture, social and family ties, romantic relationships, everything goes through networks. But as you should know, information is power. What happens when a handful of people control these networks? When they can collect and analyze a considerable amount of information about us, decide what we see or don't see, based on their own interests? In short, we would be wrong not to take them seriously, and it's an issue, particularly in Europe at a time when we regularly talk about sovereignty. But it's difficult to create social networks. It's expensive in the first place. Few investors in Europe are willing to finance the next LinkedIn, the next Twitter, or the next Youtube. And most importantly, to hope for success, you need to convince you, but also convince your friends, your family, your colleagues because otherwise, it has no interest. This is what we call the network effect: the value of the product is proportional to the number of users using it , which makes the arrival of new players almost impossible. And yet, there is today another way . What if, instead of trying to build the next Twitter, we completely changed the way social networks work? What if we returned, in a sense, to the origins of the Web: a more open and decentralized Web, where you could choose your provider without losing access to the rest of the network? In short, let me introduce you to a protocol that could help us redefine our social networks and create opportunities, particularly in Europe: ATProto. You may know Bluesky ? Bluesky is a new Twitter with around 43 million users. If its success remains modest compared to its predecessor, its growth has been rather dynamic over 4 years and varies depending on regular waves of departures from X. Thanks Elon… But you might say, Bluesky is just another US company, so another monopoly in the making. Yes, certainly, but you need to look under the hood. Bluesky is based on a protocol: AT Protocol. This protocol defines the data present on the network, how we store it, how we view it, how we can moderate it (label it), in short, how the network works. More precisely, it means that Bluesky is not required to own your data to display it. On Twitter, Substack or Medium, the data is owned by these companies and can be filtered, modified, or monetized without your consent. Here it's different. Your data is hosted on a PDS (Personal Data Server), and this PDS doesn't necessarily belong to Bluesky. For example, my data is hosted by Eurosky , a European PDS. And I'd like us to pause for a moment on the Eurosky homepage: Eurosky is also behind mu.social , an alternative Bluesky client, and mu is the first prefiguring thousands of other social apps. Because yes, the ATProto protocol is so open that each piece of the puzzle can be replaced or extended. And that's where we can start having fun. Everyone can build on top of it . We can define our data, so we can build alternatives to Instagram ( flashes , grain ), to Pinterest ( current.is ), to Substack and Medium ( Leaflet , Writizzy ), to Tiktok ( Spark , Skylight ), to InoReader ( Gleen ), or even to Github ( Tangled ), all based on the same architecture, but especially, already benefiting from a user base of 45 million people. The network effect is already there . Bluesky is not ATProto, it's one piece of a much larger ecosystem: the atmosphere . By default, the data present on ATProto is made to display short messages. It's the famous Bluesky message you already know: a message, an author, a number of likes etc… But, at first, you can't post long messages and besides it's not really designed for reading blog articles, videos, or source code. But we can extend all of this via lexicons . A lexicon is a common language, a data schema, it's roughly the set of vocabulary on a network record, and that certain applications will use. For example, the lexicon of a Bluesky post is defined by and we'll find fields like , , , (when there's an image) etc… But anyone can define a new lexicon. Obviously it won't be read by all applications. Bluesky won't read your lexicon but your application can. And precisely not long ago I came across a lexicon that particularly interested me for blogging platforms: standard.site . The initiative comes from three platforms that partnered together: Leaflet, pckt.blog and Offprint and they proposed this new format for long-form publications. This format pleased so many that it is now available via a plugin in the WordPress ecosystem, for static blog generators via Sequoia , and now it's also becoming central in Writizzy , the product I'm building and which runs this blog (and probably soon also Bloggrify which I also maintain). This integration therefore allows you to have native understanding of your Writizzy publication in Bluesky/muSocial. Note the "view publication" link which doesn't normally appear for a simple base link. But most importantly, this allows all posts published on Writizzy to be visible in all readers that scan the atmosphere for blog articles: standard-reader , Heron , potatonet, docs.surf , leaflet etc… In short, in one step, Writizzy becomes a member of the atmosphere and can expose a new article to millions of users. And this example shows us that we can now create new apps, with their own vocabulary, on existing PDSs and benefiting from an already established network. Beyond this specific example, ATProto offers many opportunities to return to a more decentralized web under your control. You could, for example, have your own PDS (personal data server) that participates in the AT Proto network. And you can go even further. The entire protocol provides that each feature is composable and that includes moderation or feeds. It's not Bluesky that imposes them on you. Again, Bluesky provides a service, but it's not ATProto. If you want to use another labeler (a tool that allows moderation), or if you want to create your own feed with your topics, you can. I could create in the future a feed of all Writizzy tech blogs for example. In short, you can regain control over your data, the algorithms that push content to you, and the associated moderation. But more than that, it's huge opportunities to rebuild all the social networks we're missing in Europe: LinkedIn, Youtube, Tiktok, Instagram to name a few. Hoping that indeed, mu.social is only the first among thousands. This open and composable web won't happen by itself, and it won't come from Silicon Valley. Today, the infrastructure is ready, the user network exists, and the foundational building blocks are in place. It's up to us, European developers, creators, and entrepreneurs, to seize ATProto to build tomorrow's platforms, on our own terms. PS: oh yes, and if you comment below the post on Bluesky/mu.social, the thread should come back as a comment on the blog

0 views

What Is Reasoning

A few weeks ago a paper was shared that showed how to extract reasoning traces from closed-weight models. Together with online discussions about tricking models into leaking them, it made me investigate it more out of curiosity. Twitter seems full of half-truths and confusion about how this works, so perhaps this helps some to understand what is happening. Reasoning traces are usually hidden from us. We have lamented this , but mostly have to accept it. Open-weight models thankfully reveal them, and from their behavior you can see that their traces can be long and confusing. This is probably a good reason to separate them from what is normally shown to users. At minimum, UIs need to detect them. The industry has done a good job at making reasoning traces sound special and exotic, but they really are just text: the model is trained to emit its thinking into a scratchpad as part of its response, before its final answer. GPT-OSS’s Harmony response format makes this easy to see: The markers are special tokens, but the reasoning between them uses “the same text” as the final answer (just that GPT chain-of-thought text sounds really funny). When the model samples the channel token, a parser routes the following text into a separate stream exposed through the Responses API. For closed models, presumably a simple model redacts and summarizes it. How much budget goes to reasoning? Earlier APIs exposed reasoning token budgets, making it seem like a property of the sampling process. In reality, reasoning effort is baked into the system prompt. GPT-OSS puts this into the system prompt: That’s it. Training produces the resulting behavior, such as emitting the token sequence that switches to the channel. This also explains why changing the effort invalidates the KV cache. I think closed GPT models call reasoning effort “juice,” since you can ask most models how much juice they have. In DwarfStar for DeepSeek with max reasoning this is added to the system prompt: The destination of reasoning tokens is therefore a learned convention: the model is trained to keep scratch work out of the channel. Trick it into thinking it is in that channel and it may leak tokens. We have even seen older models, when thinking is disabled, reason into the bash tool and echo their thoughts to . So in some sense the only “special” behavior for some models is not to think. That at times is done by “mechanically” removing the model’s usual ways to think. In DwarfStar , disabled thinking uses the prefill , while enabled thinking uses , which are the tokens that close and start thinking. GPT-OSS doesn’t prefill but lets the model decide either way on its own. But presumably, some inference APIs prefill the opening token when reasoning is enabled, so the model never samples it itself and might prevent the sampling of the reasoning token when disabled since it can be trivially detected. This may explain why a custom tool can trick models into putting some reasoning where it should not go — but only when native reasoning is disabled. Hilariously enough I was unable to use GPT 5.6 terra for spell and grammar checking on this blog post because of safety filters. Had to switch to Kimi.

0 views
Lalit Maganti Yesterday

Opus 5 doesn't use em-dashes in code comments

I’ve been using Opus 5 as my main coding for the last few days 1 , and I noticed something odd: the code comments stopped using em-dashes; everything is double hyphens ( ) now. Instead of going off a gut feeling, I decided to actually run some checks. agentsview indexes all my conversations into a SQLite database, so this kind of question is just a query away. Here’s the chart: Seems pretty clear to me! But you might say “maybe it’s not the model, maybe it’s something else”. Well actually the cleanest evidence came from a session where I switched from Fable to Opus 5 mid-conversation and didn’t touch anything else. Looking at just that session, Fable produced 11 em-dashes in comments. Opus 5 produced 36 double hyphens. I didn’t touch anything else. What I find genuinely strange is that this only affects code comments. My chat replies still have just as many em-dashes as ever. So maybe something was tweaked when training for code-generation in particular? I don’t like Opus 5 but I ran out of tokens for Sol, Deepseek raised their prices and Fable is just too slow (and I don’t particuarly like it either!)…  ↩︎ I don’t like Opus 5 but I ran out of tokens for Sol, Deepseek raised their prices and Fable is just too slow (and I don’t particuarly like it either!)…  ↩︎

0 views

Codeberg Pages 2: Setting Up Subdomain for git-pages Back-end

Read on the website: Codeberg Pages moved to a new back-end (git-pages.) This caused me some pain. But I’m now on the other side of it, and I want to share some gotchas.

0 views
Kev Quirk 2 days ago

The Smell of AI

by Kevin Wammer Kevin talks about his opinions for various use cases for AI and LLMs. Read post ➡ I started reading this post and thought to myself "oh here we go, another piece about how awful AI is..." but as I got further into it, I found myself nodding along. While I don't feel as strongly as him about AI generated images I do agree with the premise of everything he said. It's refreshing to see some pragmatic opinions on AI and its uses as a tool. Ended up being a good read. Thanks for reading this post via RSS. RSS is ace, and so are you. ❤️ You can reply to this post by email , or leave a comment .

0 views