Latest Posts (20 found)

VLM Enhanced Metadata For My Icon Galleries

Confession: I got nerd-sniped by Sam Henri Gold’s request for my icon galleries: I'd like to humbly request artwork-level searching in macosicongallery.com What follows is a train-of-thought blog post as I play with what an implementation might look like. I’ve actually long-wanted something like this, e.g. let me search for “coffee” and show me all icons that have some depiction of coffee in them. Similarly, I’ve wanted some kind of “related” representation for icons. I have this today via existing metadata, e.g. “Show me other icons in the category ‘Productivity’” or “Show me other icons tagged as ‘orange’”. But I’ve wanted a more robust representation of this, so if you were looking at an icon that had a microphone in it, the site would say “Here are other icons that also have microphones in them.” And the relationship would be rich/smart enough to know that “microphone” was meant broadly, i.e. dynamic mics, condenser mics, ribbon mics, etc. So how would you do this? I could go through every icon one-by-one and classify/tag any attribute of its design that comes to mind, but that would take ages! Seems like a good use case for a vision model. First, I’ll look at Sam’s suggestion: run every image through CLIP. I’m not familiar with CLIP so I start with a little research: What is it? How would I use it? And most importantly: is it free/open (because I ain’t spending a ton of money to send my thousands of icon PNGs to an AI provider via their API)? Ok, so CLIP will take an image and spit back an embedding (basically a bunch of numbers representing features of the image). When you do it with multiple images, you can then compare those embeddings to see what the model considers similar (and, if you like, set a threshold for what constitutes a “match”). After getting a sense of the task in front of me, I work with the LLM to come up with a proof of concept. I don’t need to fit this into my existing site. I just want to make one-off HTML pages where I can feel out, “Can this process create anything useful? What’s the amount of work required?” This is enough to create a single HTML file where I can click on an icon and see other icons that look like it. However, I realize quickly that I’ll need to process my entire icon library to really get a good sense for how well these are matching. So I do that. [Computer goes brrrr…] Ok, now when I click on an icon that looks like a camera, I see other icons that look like cameras. Or if I click on an icon that has a checkmark in it, I see other icons with checkmarks in them — sort-of. But the results aren’t that great unless an icon is visually distinctive. I share some thoughts with Sam. He has a few other suggestions I follow. Sam mentions SigLIP2 so I start with that as a keyword. The LLM recommends DINOv2 so I say, “Let’s try it”. I give that a try, creating a separate dataset and prototype (e.g. and ) so I can continue to view these different prototypes and compare their outputs. It’s fine. Different from CLIP. Honestly not much better. So I figure let’s try another one. I go with SigLIP2. I ask the LLM to create a page where I can compare the results. Seems like six of one, half dozen of another. One does better on some kinds of icons, worse on others. The LLM recommends that, at this point, I be done shopping models. They’re roughly the same class of tool with different tradeoffs. None are breakthroughs. So now what? Sam recommends another approach: You could also try handing all icons over to a VLM, having it write up a description, and embedding THAT text against what people might search for. A thoroughly detailed person might’ve done this from the start, e.g. for an icon that’s a checkmark, add the keyword “checkmark” to its metadata. That would take me forever to go back through all my icons and do — a perfect task for a computer that never tires. So I give this a try. First I need a free/open VLM. After a little research I decide to try Moondream via Ollama . I have the machine go through each image and caption it, then pull out “tags” from the caption. For the Clear app icon , I get data like this: Then the LLM creates a single file where I can test icon matches by searching for tag overlaps (or choosing one of the popular ones). So, for example, on the search page I can click on “checkmark” and see all the icons with a checkmark. Or click on “fox” and see all the icons with a fox. Matches are pretty spot to be honest. But that’s a different kind of test than what I was doing with CLIP. Can I leverage tags for the same kind of “related icons” work that CLIP is doing? I get the LLM to cook up a single-page HTML file where I can compare “click on this icon and find other icons like it” where I’m using embeddings from CLIP vs. matching on keywords. The results seem to fare much better for CLIP. For example, here I matched on what I think of as a “checkmark icon”. You can see the approach that matches on tags didn’t work too great. I believe this is because with my simple tag-overlap approach, a distinctive keyword like “checkmark” gets diluted amongst generic tags like “square”, “blue”, and “simple”. Whereas with CLIP, if you click on an icon with a checkmark, you get other checkmarks (and not other icons that also have related tags like “square”, “blue” and “simple”). Which all makes sense. Pushing on the implementation here could help, but that’s separate work to do. I’m not sure. While doing all of this was an interesting technical exercise, there are a few important considerations I need to think through before implementing anything, such as: I’m very picky about adding new dependencies to these icon projects. I like to think that’s why I’ve been able to maintain and continue contributing to them after so many years — because I make it easy on myself (good job, Past Jim). So, for example, if I make a VLM a dependency of this project such that every time I add a new icon I have to run it through to create the embeddings, that’s a big dependency cost IMO. I’m not sure I want to do that. That said, Apple now ships foundation models in macOS 27 available through the CLI (go ahead, try typing in your Terminal if you’re on Golden Gate). So if my Mac continues to be the primary machine where I add/update metadata for my icon projects, using would be a really easy/low-cost way to process each new icon to generate a caption and keywords for matching in search. But again, I don’t know if I want to do that. I wrote this post to try and work through what I want to do, but I am still undecided. So I guess the only thing for me to do at this point is hit “Publish” on this post and keep simmering on a decision. Reply via: Email · Mastodon · Bluesky Write a script that runs a sampling of icons through CLIP’s image encoder Read the file locally, e.g. 256x256 pixel icons seem to be enough, as the CLIP model I’m using preprocesses them to ~224px anyway. Create a dataset representing the “embeddings” (an array of numbers) for each icon that I get from CLIP, e.g. Create a dataset representing the top matches between different embeddings, e.g. Create a file has both datasets (plus supplementary icon metadata I already have), render all the sampled icons, and support an for each icon that shows the icons. CLIP: click on an icon and see other icons like it. Tags: click on a keyword and see other icons with that keyword. What kind of functionality do I actually want? A “related icons” feature? Does it match on keywords or embeddings? A “search” feature that matches on keywords? How do I build these features into my codebase now, given the thousands of icons that already exist? How do I maintain this feature in the future? e.g. every time I add a new icon to my gallery, is a VLM now a dependency of this project? Given all the above, what’s the time and money cost?

0 views

2026-09-29 16:03: I think it's time... #Ubuntu

I think it's time... #Ubuntu Thanks for reading this post via RSS. RSS is ace, and so are you. ❤️ You can reply to this post by email , or leave a comment .

0 views
Unsung Today

Book review: What Can A Body Do?

A new confession: I struggle with a lot of accessibility writing. I am not necessarily saying this to blame people writing in this space, whose work I respect even if it’s not resonating with me. I think it’s more of a recognition of how difficult it is to talk about accessibility and disability. Advocacy is never easy. It’s hard not to be impatient with the pace of progress, hard to contain your incredulity that others don’t get that what is so obvious to you, hard not to constantly expect more from people, hard to imagine others – and the society as a whole – not caring. On top of that, accessibility is a more difficult subject than many I cover regularly on this blog. There are hard-to-explain technical issues wrapped in deep legal considerations, and some aspects of disability are taboo or taboo-adjacent. To me, all this often makes even best-intentioned accessibility writing less approachable. I am also, personally, to blame. I sometimes struggle with design writing, too. Some of the taboos and hangups live inside me. I might be obstinate when corrected. I have also personally seen a few times the opposite of the curb-cut effect : not fully thought through accessibility accomodations making software worse for some groups of people. All of this is making me a tricky customer. I’ve been wanting to find accessibility writing that would work for me, with just the right amount of push and pull: detailed but not overwhelming, inspiring but not cheesy, gentle but not basic, pleasant without being cheap. This book by Sara Hendren is as close as I’ve ever been to all this. It’s exceedingly well-written, to a point that it’s really hard not to quote it excessively. Witness the very opening: Every day every body is at odds with the built environment. Bodies come up against stairs and sinks and subway platforms, sometimes with ease and grace and sometimes blundering and awkward, over hurdles, even in a sudden clash. Each flesh envelope is miraculous and mundane in tis way, lugging all its gear and getting where it needs to go. Maybe you handle a sharp knife with enjoyment of its grip; maybe you wince as you sit down in or get up from your office chair. Or, just a few pages in: But consider: this dual job that design has to do is a mammoth task! How do you make a charismatic thing – not just a thing that works, but a thing that has elegant presence or pleasure in its handling, some kind of draw, a thing that pulls you in or makes you think while also being handy, modest, even garden-variety in its value? This combination is what makes design so interesting to so many people. It’s not just the quest for a better mousetrap, and it’s not just a free-form experiment, and it’s not just a slick new color scheme. […] When Hendren early on decides to focus on the right words and the best definitions – a theme permeating the entire book, important in an area of sensitivity around language – it has none of the awkwardness of a wedding toast’s “the dictionary defines X as…” but is instead a lucid exploration of the power and beauty of language, piercing right through some of the taboos I mentioned: Among disabled people there’s a bigger catch-all term, a slang for this particular mismatch: it’s called life on “crip time.” Crip is short for cripple, a name that disabled people have repurposed in an act of political reincarnation, dropping the degradation attached to a word that was used to describe them in the past, cripple , and investing it with in-group pride. “Crip time” is flexible shorthand in disability culture, used to indicate a range of uneasy relationships to thje pace of contemporary industrialized life, with its relentless and clock-driven organization of hours and days. As a disabled person, to say you’re “on crip time” on a Tuesday might signal te extra time that it takes you to get to the train platform or in and out of a public bathroom. It can also stand for bigger systemic fits and starts – the spiky, unpredictable time it might take a person to proceed through a fairly rigid K-12 education that’s built on all kinds of normative chronologies. This whole book is this – it’s accessibility and disability as details, as creativity, as confidence, as joy, as warmth, and as craft. The writing has so much range and color, covering some of the aspects I expected, but also talking about autism, dementia, phobias, and general intellectual disability (who knew? your brain is your body, too). Hendren dips into history as needed – the origins of the usually simplistically portrayed curb-cuts is fascinating – but the book is anchored in the present. There isn’t a lot of software in the book, and the other reason I’m giving it only four stars is that I wanted… more photos. I’m a single-issue voter this way, always dreaming of visual evidence of devices, people, and places. But don’t let this stop you. This book will help you understand accessibility and disability – and if you already do, it might help find good ways to talk about it. You should read it. = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/book-review-what-can-a-body-do/1.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/book-review-what-can-a-body-do/1.1600w.avif" type="image/avif"> (Thank you to Ethan Marcotte for suggesting the book and a conversation about it, and to everyone who answered my question about good accessibility writing – check out the answers on Mastodon or Bluesky if you are interested in more recommendations!)

0 views

2026-09-29 11:10: Soooooo Hacker News happened. 🤣

Soooooo Hacker News happened. 🤣 Thanks for reading this post via RSS. RSS is ace, and so are you. ❤️ You can reply to this post by email , or leave a comment .

0 views

One More Note on Agents, Meta Connect, Meta Enterprise Platform

Meta has the chance to own the consumer agentic space; going for enterprise is a big mistake.

0 views

October Is Spooky Video Game Month

It’s just two days shy of being October, the spookiest month of the year. October does not just signify Inktober or any other compound word involving the name of the month. No; October means playing Mega Ran, Richie Branson & Kadesh Flow’s Ghouls ‘N Ghosts albums on repeat ( we’re the kings of the night; GnG resurrected ). I have high hopes of them releasing album number five, but in the meantime, just listen to a few tracks of the latest album and you’ll know why it’s part of the October canon: The first song on the album is a rearrangement of a classic Castlevania track—something Mega Ran loves to do and I love to listen to. The lyrics of the linked track—nr. 2; Make It Out (Siren Head) —set the tone of a dark haunted castle that, once entered, refuses to let you go: Too much time trying to get away Eerie sounds almost had me swayed Thinking there was someone to save Got me dodging an early grave […] I don’t know if I’ll make it out (4x) But Lord knows I’m trying Speaking of which; why don’t we all play as many Castlevania games as we can in October? We’ve got two things to celebrate: the spooky season (and thus the harvest of ample of pumpkins I guess) and the arrival of Castlevania: Belmont’s Curse , a new 2D entry in the mainline series! If you don’t count Igarashi’s Bloodstained series as a mainline Castlevania game, it’s been since Order of Ecclesia in 2008—18 years ago—that we ‘Vania fans are being treated to a brand new 2D metroidvania game. I’m not counting 2009’s The Adventure ReBirth ; the remake of the 1989 Game Boy game; nor 2024’s Haunted Castle Revisited ; the remake of the 1987 arcade game. Yet perhaps it is worth it to paste in a flyer for the arcade game to help get you into a creepy mood. Haunted Castle: Re-Vamp An Old Game With This New Kit! 1 To 2 Players. Doesn’t that penetrating gaze from the beautifully dressed woman kneeling next to a grave marked as “STONE” without getting a single spot of dirt on her very visible shiny legs send a chill down your spine? Or send anything down anything else? How about the gaze of Dracula himself, staring dangerously looking into the foggy distance, or rather, having this I-need-to-pee look? That was the late eighties—they don’t make them anymore like this. But no worries: they are still making a Castlevania game! Someone at Konami finally resurrected the genre after realising that this could be a sure way to fill the coffins back up with lots of cold hard cash; needed for research purposes to do execute some more cross-breeding vampire-werewolf-human experiments. I don’t understand why they realised it so late but better late than never. The Castlevania Collections on multiple platforms are still selling well; shouldn’t that be an indication for an appetite for more? In any case, I will enter October prepared. I already went ahead and played a couple of 16-bit entries that I somehow missed, such as Bloodlines on the MegaDrive and Vampire’s Kiss on the SNES. I play Aria of Sorrow each year: to me, it’s simply the best game in the series, and one of the best games ever created. It executes every typical ‘Vania mechanic flawlessly, has a soundtrack that just slaps, and doesn’t outstay its welcome. Obviously, it’s in my Top 20 GOAT list . Yet for being such a fan of the more “modern” metroidvania adaptations the series evolved towards after Rondo of Blood , I have never truly finished thé most important entry: Symphony of the Night . I wasn’t a Sony/PlayStation kind of person and even after finally having access to it through a hidden unlockable in the Dracula X Chronicles remake on the PSP, I had other more pressing things to do such as trying to figure out how this thing called working works. The game arrived in Europe a year after I graduated. Each year in October, we at the DOS Game Club play a spooky DOS game . Last year, that was Halloween Harry . In two days, you can join us playing Legacy: Realm of Terror . In Legacy , you explore yet another spooky mansion—this time not Dracula’s, but from someone or something much worse: something that serves as a gateway to ancient eldritch horrors. Why don’t you create your own spooky video game backlog to try and work through in the upcoming month? Here’s mine: I don’t expect to finish them all but if you are interested in exploring the Castlevania series, I do recommend to just start with the beginning. The original “classicvania” stage-based games are very short—but also very hard: that was the accepted way to artificially prolong the game length in the eighties/nineties. Play them through the Anniversary Collection or en emulator that features save states in case you get frustrated by bats knocking you off staircases. Trust me, you will. There are enough “best ‘Vania games” lists floating around on the internet. For me, the GBA and DS games are firmly placed at the top. Skip a game if it doesn’t click with you in thirty minutes. Enjoy whipping and be whipped by Death! And don’t forget to make good use of the bathroom if you suddenly find yourself staring at the distance. Related topics: / games / By Wouter Groeneveld on 29 September 2026.  Reply via email . Castlevania: Dracula X Chronicles (the Rondo of Blood PSP remake) Castlevania: Symphony of the Night Diablo II: Resurrected (I’ve played the original to death and feel the sudden urge to do a few more Mephisto hell runs. Plus, the new class looks amazing: official new content in a game that’s basically 26 years old!)

0 views

Donating to open source

I have decided to start donating a small amount of my fun money to open source projects. After some research, the initial set of recipients are These are important to me, they need more money, and I believe they can use donations well. If you disagree, I’d genuinely like to hear what mistakes I’ve made in the analysis below. I hope it shows that this is important to me. (Continue reading the full article on the web.) Perl 5 Core Maintenance Fund

0 views

Dutch Police Arrest ‘Reformed’ Hacker in Shiny Hunters Investigation

Authorities in the Netherlands have arrested a 24-year-old convicted cybercriminal on suspicion of aiding in data thefts and extortions by the prolific hacker group ShinyHunters . In the days immediately following the suspect’s arrest, remaining ShinyHunters members dramatically escalated their attacks, stealing highly sensitive data from the FBI and extorting the Russian ransomware group Cl0p . According to three sources familiar with the matter, the Dutch man arrested by authorities this month is Pepijn van der Stap , a convicted cybercriminal from Almere and Lelystad in the Netherlands. Van der Stap was previously convicted in 2023 in connection with a string of data thefts and extortions that prosecutors said earned between €1.5 million and €2.7 million. At his trial in late 2023, van der Stap admitted that he lived a Dr. Jekyll and Mr. Hyde existence, secretly using the hacker handle “ Umbreon ” to extort victims and post their data on English language hacking communities like the now-defunct RaidForums and Breached. By day, however, van der Stap was working as a software engineer at the Amsterdam-based cybersecurity startup Hadrian , while volunteering at the Dutch Institute for Vulnerability Disclosure (DIVD), a nonprofit security research group. Pepijn van der Stap’s alter ego “Umbreon” selling a database on RaidForums, offering information on 2.3 million people from The Netherlands in September 2021. This user’s avatar is a depiction of the Pokemon character Umbreon. Image: KELA. Van der Stap confessed to his data theft and extortion activity, and was sentenced to four years in prison (one of which was suspended). During his trial, van der Stap opted to remain in custody for a time rather than at home, saying he could not find better treatment on the outside for his ongoing psychological issues, which he claimed included PTSD related to childhood trauma. He was released from prison in December 2025. In an interview with KrebsOnSecurity on September 9, 2026, Van der Stap cast himself as a reformed hacker who was trying to turn his life around and make a positive contribution to society. Van der Stap is currently employed as offensive security lead at the Dutch company Neo Security , which did not respond to requests for comment. Van der Stap said he was still dealing with civil lawsuits and restitution related to his previous cybercrime victims, and that he was trying his best to make amends. But not long after that interview, the Dutch hacker abruptly stopped replying to messages. Efforts by others close to him also repeatedly failed to elicit a response for the past two weeks. The LinkedIn profile for Pepijn van der Stap. According to two sources with knowledge of the matter, Van der Stap was arrested by Dutch authorities on or around September 16, and has been held in custody for questioning since. One source said a colleague of theirs personally witnessed Dutch authorities carting items out of Van der Stap’s residence. Authorities in the Netherlands have been asking the public for help in identifying the voice in a recorded telephone call from February 2026 in which a native Dutch-speaking ShinyHunters member social engineered their way into Odido , the nation’s largest mobile telecommunications provider. In that intrusion, ShinyHunters tricked an Odido employee into logging in at a spoofed website, and then used that access to steal data on more than 6.2 million Dutch people. Responding to Dutch news media, ShinyHunters confirmed that the suspect in the audio clip is indeed a member of the hacker collective. “Our team member has our full support – emotionally, mentally, and financially,” the hackers said. “Everything has been arranged, including a criminal defense lawyer. We do not look down on our staff and members; we take excellent care of them,” reads a statement ShinyHunters shared with NL Times . It remains unclear if the Dutch police have matched the Odido caller to a confirmed real-life identity. The Dutch police unit handling the Odido incident did not respond to requests for comment. The group also lashed out at the authorities in the Netherlands. “The Dutch police will need all the luck in the world – and everyone’s prayers – if they want to catch him before we carry out another large-scale data theft in the Netherlands,” the ShinyHunters statement said. “Frankly, the Dutch police are a big joke; they are incapable of doing anything. Incompetent. Irrelevant. Unimportant. Useless.” Just days after sources say Van der Stap was detained by Dutch authorities, ShinyHunters claimed credit for an unusually brazen breach at the FBI’s job application site apply.fbijobs.gov. According to reporting from 404 Media , the data stolen from the FBI site includes Social Security numbers and personal information on more than 5,000 officials. 404 Media and Reuters reported the FBI data included each person’s job title or team, such as special agent, threat intake examiner, major cybercrimes unit, and those investigating cyber threats from foreign state-backed actors. Reuters examined documents shared by ShinyHunters and found they included sensitive psychiatric and medical files of FBI staff. The FBI issued a brief statement confirming the hack. ShinyHunters said it gained access to the FBI site and other victims by exploiting a recently patched vulnerability (CVE-2026-35273) in PeopleSoft , a software-as-a-service platform from the software giant Oracle that is broadly used by companies to manage hiring and human resources, benefits and payroll. Oracle quickly issued a fix for the Peoplesoft vulnerability that ShinyHunters reportedly began exploiting as a zero-day in June, and at the time Mandiant released web application firewall rules intended for organizations who couldn’t apply the security update quickly enough. But on Friday, BleepingComputer reported that ShinyHunters used a URL-encoding trick to bypass Mandiant’s suggested web application firewall rules designed to mitigate the threat from the PeopleSoft flaw. In a report released Sept. 25, security experts at Mandiant and the Google Threat Intelligence Group (GTIG) confirmed that ShinyHunters had mass-exploited the PeopleSoft vulnerability to steal data from dozens of systems across a range of industries, including higher education, technology, healthcare, agriculture, transportation and government. Van der Stap’s former hacker alias Umbreon was hidden in plain sight throughout the imagery ShinyHunters used to spread news about the FBI hack: The defacement image that ShinyHunters left behind on the hacked FBI jobs site included an ASCII art design featuring the Pokemon character Umbreon. The message at the top read, “This site has been seized by ShinyHunters. rooting your systems since ’19 ;)” The image appears identical to a defacement message ShinyHunters used in their 2020 hack of the English-language cybercrime community Hackforums. The defacement message left by ShinyHunters on the FBI jobs site included an ASCII art rendition of the Pokemon character Umbreon. Image: Bleeping Computer. Multiple sources close to the ShinyHunters investigation said the group’s recent risky attacks against the FBI and one of Russia’s most venerated ransomware groups amounted to a major pivot away from the more measured tenor of the hacking gang’s operations. Those sources said the sudden shift came about after ShinyHunters was taken over by a teenage cybercriminal from Amman, Jordan who goes by the nickname Rey and operates as part of a cybercrime group called ScatteredLapsussHunters (SLSH), which experts say is an amalgamation of three hacking groups — Scattered Spider , LAPSUS$ and ShinyHunters . Those sources said Rey had an ongoing beef with the Dutch hacker over control of the ShinyHunters brand and data, and that the inclusion of the oversized Umbreon Pokemon image in the FBI jobs site defacement was likely an attempt by Rey to pin the hack on the Dutchman. Rey was first publicly identified by the cybersecurity firm KELA in March 2025. In advance of our November 2025 profile of Rey , KrebsOnSecurity messaged Rey’s father and asked for permission to interview his teenage son. Rey’s dad merely forwarded the message to his son, who admitted to participating in ransomware attacks and said he was trying to extricate himself from the SLSH hacker group. Immediately after news of the FBI jobs site hack was picked up in the media, Rey’s main account on Twitter/X (Ryan Moran/@rmoskovy) was taunting the Cl0p ransomware group and the FBI, crudely depicting them as the twin towers in New York being struck by planes labeled “cl0p drama” and “fbi breach claim.” In the foreground of the city is the giant Pokemon figure of Umbreon. A taunting meme uploaded to Twitter/X by Rey’s now-defunct account on Sept. 22. A giant float-sized version of the Pokemon character Umbreon can be seen in the bottom left. On Sept. 24, KrebsOnSecurity again contacted Rey’s dad, asking to interview him and his son for a story on Rey’s apparent ascendency as the head of ShinyHunters. Just hours after that request, Rey deleted his longtime Twitter/X account. Meanwhile, Rey’s dad, who works for the Royal Jordanian Airlines, has failed to respond to a half-dozen emailed requests for comment about his son’s alleged activities. Where does the bad blood between SLSH and ShinyHunters come from? According to a story in Wired this month, ShinyHunters and SLSH members briefly partnered earlier this year to help better monetize important stolen credentials collected by TeamPCP , an upstart group that was having great success compromising global code supply chains with malicious software but hadn’t been able to profit much from their stolen data (two alleged leaders of TeamPCP were arrested last month in Australia , and in an interview the TeamPCP leader claimed they made just $20,000). The Wired story noted how Mandiant had infiltrated TeamPCP and was secretly responsible for having the crime group’s stolen credentials burned so quickly: Mandiant was secretly feeding those credentials to the major cloud providers like Amazon and Microsoft, who quickly invalidated the stolen keys. Meanwhile, the formerly cooperating hacker groups began to blame one another for causing the credentials to become worthless. Wired’s Andy Greenberg reported that a few weeks after partnering with TeamPCP, “ShinyHunters went rogue, carrying out its own extortions with TeamPCP’s credentials but without giving the supply-chain hackers their cut.” Mandiant researcher Austin Larsen told KrebsOnSecurity earlier this month that ShinyHunters has been enjoying a successful extortion spree so far this year, and is on track to pull in nearly $100 million in extortion payments from cybercrime victims in 2026. Van der Stap claims he was never motivated by money and that his earlier hacker activity was driven by a desire to have the world’s most complete collection of stolen databases. Speaking with reporters from Bloomberg in 2024, Van der Stap said that singular focus in turn fueled his desire to carry out cyberattacks. “The hacking was very easy for me, and it wasn’t a compulsion,” he told Bloomberg . “My habit was collecting. Collecting data, organizing data, downloading data, creating folders.” DIVD, the nonprofit security research group where Van der Stap previously served as a volunteer, disclosed on LinkedIn last week that the organization was dealing with an internal cybersecurity incident that appears to have involved the malicious use of artificial intelligence. DIVD has released few details about that incident, but a spokesperson for the nonprofit told KrebsOnSecurity it does not appear related to ShinyHunters, nor are there any signs the matter involves the work of a previous volunteer. Update: 3:44 p.m. ET: Corrected Van der Stap’s age, which is 24 (not 23). Update, 4:54 p.m. ET: The Dutch police have confirmed the arrest of a 24-year-old in connection with the ShinyHunters investigation. In a statement on Twitter/X , the Dutch police said the man will appear on Tuesday, September 29 before the chambers of the Rotterdam District Court, and that it will provide more information tomorrow.

0 views
David Bushell Yesterday

Shin honkaku

Shin honkaku is a sub-genre of Japanese detection fiction — and my new obsession! I’ve even designed my own book cover, scroll down to see… “Shin honkaku” translates to “new orthodox” and pays homage to classic western authors of the early 20th century such as Poe, Queen, Van Dine, Carr, Christie, and Conan Doyle. All of whom are often referenced by the fictional detectives within shin honkaku books. The genre began in the 1980’s. As the name implies, it’s a return to the roots of mystery fiction. There is a deliberate emphasis on “fair play” allowing the reader every chance to solve the mystery. Prior to shin honkaku, Japanese detective fiction had a seedy reputation. Sōji Shimada writes in the foreword to The Moai Island Puzzle : In the 1920s, when Edogawa Rampo imported the Edgar Allan Poe-Arthur Conan Doyle style of mystery novel which started in the West, and new Japanese writers gathered to solidify the form of this new genre, the scientific revolution which had given birth to the detective novel had not yet arrived in Japan. In the absence of scientific inspiration, Rampo turned to the grotesque haunted house attractions of the Edo period (1603-1868). Eventually, Rampo’s followers started going too far, even introducing the pornographic tendencies of Edo period entertainment fiction into their novels. Foreword to The Moai Island Puzzle (Alice Arisugawa, 2016) - Sōji Shimada This effectively killed mainstream interest in mystery fiction, as Shimada continues: As a consequence, Japanese mystery authors were looked at with contempt by authors of pure literature at the time, purely because of the vulgarity which offended their morals. Hence the belief that mystery fiction was just lowbrow entertainment found its way into Japan’s literary world, as well as with the readers. It is Sōji Shimada himself who is often credited with the revival. Shimada’s debut novel The Tokyo Zodiac Murders (1981) is considered one of the best. It’s certainly my favourite! Somewhat ironically, it was shortlisted for the Edogawa Rampo Prize. This book was where I discovered the shin honkaku genre. If you follow my notes blog you’ll have read my short summary: The story itself is terrifyingly gruesome. The mystery behind who committed the murders is a seemingly impossible series of events with a deceptively original explanation. After re-reading the details more than once I “solved” the mystery. At least, I got the gist of how things happened and arrived at the correct name(s). The exact explanation is far more elegant than I pieced together myself. Note for Mon 31 Aug 2026 - David Bushell Since reading Sōji Shimada’s two translated books, I’ve read Yukito Ayatsuji , Alice Arisugawa , Keigo Higashino , and Seishi Yokomizo (older honkaku, but worthwhile.) The stories are short relative to my usual epic sci-fi/fantasy tomes. I suspect my perspective on typical novel length has been warped. If you’re a fan of puzzle games, this genre will be your cup of tea. The books often include dramatis personae and illustrations for reference. And red herrings, of course. So many red herrings! Back when I studied design I created a series of space opera paperback covers. This project was a significant part of my graphic design bachelor’s degree. The brief was also a “competition” ( cough — exploitative spec work — cough ) that could have seen my designs become a limited print run. An insider said I was shortlisted to win, but my printing techniques were not financially viable. Although in hindsight maybe they were being kind… either way, the winning designs were very nice indeed. Inspired by my new obsession for shin honkaku, I dusted off the ol’ image editor and whipped up a design mock-up for The Tokyo Zodiac Murders . For a quick weekend project, I’m very pleased! I would love to tighten up this design and create a full series of shin honkaku books. I forget the rules and limitations of book design, maybe I’ve committed a fatal folly? Doesn’t matter! It’s a concept. Early in my career I went from graphic designer, to web designer, to front-end web specialist. It’s nice to go back to where it all began. Still got it! I would welcome book suggestions in you’re familiar with shin honkaku. I fear I may exhaust the available English translations. I used the following resources in my design: The Tokyo Zodiac Murders was written by Sōji Shimada in 1981. I read the English translation by Ross Mackenzie & Shika Mackenzie published by Pushkin Vertigo who have a great collection of designs for other crime authors. Thanks for reading! Follow me on Mastodon and Bluesky . Subscribe to my Blog and Notes or Combined feeds. Cry Wolf font by Hanoded (licensed) Hina Mincho font by Satsuyako (Google Fonts) Zodiac sign icons by Chaiconator from Noun Project (CC BY 3.0) Shoe print by rasendria from Noun Project (CC BY 3.0) Perfect-bound paperback book mockup (free)

0 views

Lavinia

Lavinia, daughter of one king and wife to another, is barely mentioned in Virgil’s Aeneid, where she is just a wisp of a girl, ripe for marriage and little else. But then in the final hours of his life, Virgil finds himself with Lavinia in the sacred woods of Albunea, and discovers she has a voice and desires all her own. Lavinia is surprised to discover that she has been breathed into life by a poet, but the clues he tells of her life are enough for her to begin to weave her own story, one that is infinitely richer than any he might have seen. As Virgil leaves her life untold and unfinished, it becomes no one’s life but her own. View this post on the web , reply via email , or become a supporter .

0 views
ava's blog Yesterday

book club: kiss of the spider woman by manuel puig

In the Grizzly Gazette book club, we've voted on reading Kiss of the Spider Woman by Manuel Puig next, suggested by Psycheoma . We kinda lost sight of the book club and posting it in time :p I've struggled with reading fiction for a while now; I manage to squeeze some in here and there, like when I read The Dispossessed 3 years ago and loved it, but most of the books I own and end up finishing are non-fiction. It wasn't always this way, especially as a child, but the older I got, the more I felt pressured to read things that had an obvious "return of investment". Fiction can teach you a lot in unexpected ways, but you never know whether it will just be a nice story for you, or something that makes you think of life, society or yourself differently, and in which ways that will be. You have to start reading and trust the process, just enjoying it for what it is, and let it give you whatever it ends up being, even when it is not life-changing or improving on anything. Meanwhile, buying a book specifically about a specific topic/field or an autobiography immediately tells you upfront what you can expect and what you'll be taught, no surprises. Having read it carries the promise of improving in that area or at least having valuable knowledge. I like that predictability and usefulness, and I am a very impatient person with a tough schedule at times, so I am unfortunately obsessed with not wasting my time. After all these years of barely any fiction reading, I've become lazy and inflexible. It's now so, so hard for me to give books a chance without knowing definitively if I'll enjoy them and what I'll take away from them. Non-fiction is easy for me to read because I already have a reason to care: Me. I'll relate it all back to me, and my interests, what I already know, or who or what I want to be. Anything that is not relevant or interesting to me can be disregarded in self help. Fiction challenges that, because the beginning of a book will just present me with characters that do not yet give me a reason to care. I don't know them, I don't know anything about them yet, and I cannot relate them back to me. I'll have to sit there patiently until everything is slowly revealed and the characters become more than random names on a page, and may be completely unlike me. That isn't easy for me, but I'd like to change. I went in completely blind. I had not looked up anything about it and had no idea what to expect. The narrative style (a dialogue without names showing who says what, but it becomes apparent when you continue reading and see how they refer to each other) is definitely a surprise and took me a while to detect and get used to, but then it is fairly easy to read and understand. I was done with it after 2 days, in which I read for about 4-5 hours in the morning. The book released in 1976 centers two prisoners in Buenos Aires in 1975, one a trans woman referred to as Molina (her last name), the other a man named Valentín. While Molina is incarcerated for "corrupting a minor", Valentín is a political prisoner because of his involvement with a revolutionary group. The story is told through their dialogue, which is mostly Molina describing movie plots to Valentín, many of them centering on romance and having real life equivalents. Through discussing the movies, they occasionally reveal things about their outside life or themselves. Later on, the story continues through written reports by a third party. Molina's gender is mostly not acknowledged; she repeatedly states she sees herself as a woman and is not a man, yet the prison system and even Valentín continue to misgender her and simply see her as a homosexual man. Even the English Wikipedia page makes no mention of it. (Spoilers ahead) After about half of the book, it gets revealed that Molina is in this cell together with Valentín because she is cooperating with the prison on getting more information out of him so they can arrest more members of the group. Her meetings with the warden are covered up by getting groceries and pretending they were brought by her mother and lawyer, and her incentive is an early pardon. While everyone involved seems to not (fully) acknowledge Molina as a woman, they are absolutely aware that she is queer, to them an effeminate gay man, and use what they associate with that femininity to their own gain. And Molina indeed fits the bill: A caring and romantic person who wants to be liked, wants to help people and build a bond. She helps Valentín as he is sick from food that was intentionally tampered with to weaken him, makes him food and tea, even washes his stuff and gives him hers when he shits himself from it. She even has sex with him, seemingly consensual. It reminded me a lot of the practice of V-coding , which is the common practice of placing trans women in the same prison cell as male inmates to placate them. The idea is that by getting to rape the trans woman, he'd become calmer and more docile. It's absolutely horrible. While in this book, Molina apparently wasn't raped and wasn't just placed there to calm down an aggressor, she nonetheless seems to have been used specifically for her femininity and willingness to have sex with men, in the hopes to break through the walls of a very closed-off man and get some intel. The warden and sergeant even toy with the idea to release Molina, orchestrate a fake admission by her and surveil her in the hopes that the revolutionary group would seek her out to get revenge, which would give them a chance to pounce. It's a great fictional, but sadly very realistic case showing how far the justice system was (and is) willing to go to abuse and endanger queer people if it helps their investigations. Molina seems to want to be a neutral party to all of this initially; telling the warden openly she's making no progress, but asking for extensions. She has become attached to Valentín, and who knows who else she could be matched up with? It could be worse. At the same time, she gets multiple chances to press him on his political contacts but chooses not to, even one time telling him that she'd like to hear anything else but the name of his comrades. She fears that they could interrogate her and she'd spill, both betraying him and painting a target on her back after her release. Slowly, she changes her mind, and Valentín convinces her to relay messages to his comrades once she is out, telling her who to call, where and how. He wants her to get involved with the group(s) too. Unfortunately, she gets surveilled for longer than Valentín said she could expect, and in ways she possibly failed to detect after a while. Thinking she was safe, she makes moves that make it obvious to surveillance that she is reaching out to the group and trying to get picked up by them. Sadly, at one final attempt (which already seemed decently hopeless and dangerous to Molina, as she already withdrew all she could from her bank account and left it to her mum), she gets nabbed by agents just as members of the group drive by to kill her and wound an agent, likely to prevent her from spilling anything. It feels tragic, another trans woman killed, used as cannon fodder, used in the agenda of the men around her, just seeking for connection, love, a bigger purpose. One can only hope she felt at peace as she died for a supposedly greater good. Meanwhile, Valentín is still in prison, getting tortured during interrogations. He thinks of both Marta, his love, and Molina, as his mind drifts off on a high dose of morphine in the infirmary, even mixing them up. He regrets what happened to Molina because of him. My wife asked me if it was a fun book. Kinda depends on the definition... I was interested in it enough that I was able to read it for hours on end. Was I truly excited or gripped, always expecting the next turn or twist soon? Mostly no, it was very predictable in the sense that the prison days would pass and another movie would be discussed; I think I was only glued to the paper when the surveillance reports came up, because I wanted to know how Molina would act - would she really follow his instructions, and what was her support network outside like? I wasn't laughing or crying either, I had no strong emotions, but it was nice to read, it kept me company, it filled my time in a way that didn't feel like a waste. It's not a book I would put on a list of recommendations, but if anyone was planning to read it or asking if it was any good, I'd encourage them to. Published 28 Sep, 2026

0 views

The life triangle

Over the past few years, I spent way more time than I probably should have thinking about my life. And the way I usually look at it is by asking two fundamental questions. The first question is “do I want to have a life?”. You might think this is a silly question to even ask, but since I was a teen I’ve always had a weird relationship with the concept of death, and as a result of that, the whole idea of being alive at all has never been an absolute certainty. Which is why I periodically ask myself the question, and so far the answer has clearly always been yes, and that is why I’m still here. Figuring out an answer to this question is often pretty straightforward, and the answer is binary, so there’s no point in discussing it further. The second, and more interesting, question I ask myself is “what kind of life do I want to live?”. Finding an answer to this one is a lot trickier. But one framework I’ve been using lately is to visualise my life as a triangle. I’ll explain what I mean, but I don’t plan to add a drawing so you’ll have to picture this in your head. The first vertex represents the “individual me” . These are all the things that are related to myself as an individual, as if I was existing in a vacuum. Things like health, personal aspirations, dreams and goals, experiences I want to live, skills I want to acquire, my job and so on. I know we don’t really exist in a vacuum, but you get the point. The second vertex is the “social me” which is the collection of things that are related to other people: relationships, responsibilities toward others, shared goals and dreams, and so on. The third and final vertex is reality. It represents the current state of my life at this specific moment in time. These three points are obviously connected; this is a triangle after all, and the distance between them represents how far my current life is from what I want and also how far apart my individual life is from my social life. Those points also define a plane that represents the infinite number of possible lives that are available out there. In a perfect world, this triangle is collapsed down to a single point, because that would mean that there is a perfect overlap between what I want for myself, what I want for my relationships with the other people in my life, and the life I’m actually living. The aspiration is to get to that point, but making the triangle as small as possible is a worthy endeavour. I find this framework helpful because if the goal is to make the overall area smaller, there are a couple of different ways I can go about achieving that. I can try to improve my life and move that vertex closer to the other two; that is one obvious approach. But I can also decide to willingly move one of the other two vertices: I can change my goals and my ambitions, for example, or I can change the people I surround myself with in order to live relationships that are more aligned with the life I want for myself. Or I can combine all those approaches; there’s plenty of flexibility here. What I do know is that lately I felt this triangle grow bigger, which is precisely what I don’t want it to do. But I also lost track of where those dots are, which is also not ideal. It’s funny how there are people who have some things figured out , and I’m sitting here, looking inward, and the more I look, the less all this mess makes any semblance of sense. But I guess that’s part of the fun. Thank you for keeping RSS alive. You're awesome. Connect via email :: Sign my guestbook :: Support for 1$/month

0 views
David Bushell Yesterday

Leaving them behind

Quick note before we begin: this is the first of two posts I’m publishing today. You’re welcome to skip ahead to: Shin honkaku — it’s far more fun! I’ve waited long enough! I’ve entertained one “wait six months” too many! The TL;DR for my updated AI policy has changed: The absolute vileness of the AI industrial complex knows no bounds. Beyond morality — because let’s be honest few care — it’s very simple: There is no worthwhile career in AI- anything . Simple as that. The AI industry is designed to dehumanise and commoditise labour. Everyone who has dedicated their life to token servitude has become a dull fungible meat proxy. The software and web development industries are leading this brain drain. I’ve observed devs go from the giddy thrills of gambling with their employer’s tokens, to the depressing realisation that they’ve been fooled by a small group of grifters and influencers. So many developers are giving up. Many have literally left the industry unable to find meaningful employment. Many more have figuratively quiet-quit. They clock in to babysit chatbots with no incentive to care about the output beyond quantity. I’m done pretending there is any hope for the AI industry to redeem itself. I’m moving on to more interesting things. Barring a monumental power shift, collapse of the industrial complex, and rise in free range grass-fed “local AI” (lol) I won’t be looking back. Wake me up if anything changes! What does that mean, practically? First and foremost I will continue to build websites for real people . I set up shop as a limited company after a decade of freelancing to bolster my commitment. I will observe the AI industrial complex cautiously from afar, but I won’t allow the bullshit I see to rage-bait me. There will be times I’m obliged to call out egregious insults to my profession . Otherwise, I’ll strive to ignore the echo chamber to protect my mental health. I am distancing myself from peers I once respected who are lost to chatbot psychosis. It is not my task to help them. I have no interest in anyone wilfully funding billionaires’ fantasies. There are new people to meet who respect humanity. I feel happier about my future now. There is no longer any lingering doubt. The perpetual tech circus may be a threat to my patience and sanity but it won’t take my career. So to immediately move on to more interesting things: my latest obsession is shin honkaku detective fiction! I’d highly recommend The Tokyo Zodiac Murders by Sōji Shimada , and The Moai Island Puzzle by Alice Arisugawa — both satisfying reads. Read part two of today’s double feature: Shin honkaku! Thanks for reading! Follow me on Mastodon and Bluesky . Subscribe to my Blog and Notes or Combined feeds.

0 views
Kev Quirk Yesterday

I Switched to Brave Browser

I've been whinging about wanting to move away from Firefox for nearly a year now. It started with Mozilla's announcement about their AI strategy and my initial thoughts . Then they doubled-down and it raised my heckles even more . Then I had a long old whinge about how fucked they seem to be. I took Vivaldi for a spin , but I couldn't get used to it. It just tries to do too much, and has way too much going on with the UI, and dark/light theme switching still doesn't work despite it apparently being fixed. So I made my peace with Firefox and stuck with it. I kept Vivaldi on my machine as a backup, and over the last 6 months or so, I've found myself having to open it more and more due to something not working in Firefox. This was far from a daily occurrence, but it happened often enough for it to become annoying. But I still couldn't use Vivaldi full time. It's worth stating that things not working isn't the fault of Firefox. It's developers not testing their web apps with Firefox, since it has such a small market share. I decided to install Brave a couple months ago just to see how things went. I had some messing around to do up front, like hiding all the crypto and AI nonsense, but once that was gone, it's been fine. Everything works how I expect, both on Ubuntu and Android. The UI is similar to Firefox, and most importantly, it doesn't try to boil the ocean. I renders webpages, then gets out of my way. I don't need it to check RSS feeds, or emails, or have toolbars on every edge of the screen. Even dark mode switching works! I haven't felt the need to open Firefox since starting this exeperiment, so I think Brave is gonna stick. Yes yes yes. Brave was founded by an absolute scumbag . But there's scumbags everywhere, in all walks of life. Contrary to what some people on the internet will say, my use of Brave doesn't mean I agree with, or support, Bendan Eich's world views. I consider his personal views to be separate from Brave. You may disagree, and that's fine. But Brave is objectively a good piece of software that does what I need it to do, and does it well. So I'm gonna continue using it for the forseable future. If Firefox get their act together and become relevant again, I may jump back, but for now it will remain my secondary browser. I do still have Vivaldi installed too. So I will continue to check in on it now and then, and if I decide they're on a par with Brave on Linux (mainly light/dark theme switching is what's missing), then I'll definitely make the switch as I feel their moral compass is far more aligned with my own. Thanks for reading this post via RSS. RSS is ace, and so are you. ❤️ You can reply to this post by email , or leave a comment .

0 views
Stratechery Yesterday

Apps, Agents, and Aggregation

Listen to this post : Three revolutionary products — you know the line. A wide-screen iPod with touch controls, a revolutionary mobile phone, and a breakthrough Internet communications device. That was how Steve Jobs introduced the iPhone: Today, we’re introducing three revolutionary products of this class. The first one is a wide-screen iPod with touch controls. The second is a revolutionary mobile phone. The third is a breakthrough Internet communications device. So, three things: a widescreen iPod with touch controls, a revolutionary mobile phone, and a breakthrough Internet communications device. An iPod. A phone. And an Internet communicator. An iPod. A phone. Are you getting it? These are not three separate devices, this is one device, and we are calling it iPhone. Today, Apple is going to reinvent the phone. I’ve linked to this snippet before , usually to note how the audience didn’t really understand what an “Internet communications device” was, even though that was the iPhone’s most revolutionary capability. The phrase I’m thinking of right now, however, is “Are you getting it?” Are you getting that messaging — chatbots now, natural interfaces later — is how we will communicate with AI? Are you getting that not only will we not program computers, we won’t use them — AI will? Are you getting that pre-built UI — write once, run everywhere, for everyone — is dead? These are not three separate predictions: this is reality, right now, in 2026. The future is here, even if it’s not widely distributed. Or is it? I wrote on February 18, 2014 that Messaging was Mobile’s Killer App ; I specify the date because I fortuitously published that Article the day before Facebook acquired WhatsApp . My thesis was that phones may have started out as convenient information devices, but the killer use case was a very human one: Seven years ago, the computer became pocketable, but the original use cases were about making the passive presentation of information accessible not just at a time convenient to the viewer, but also at any place: the web was now everywhere. Still, it’s only recently that the killer app for this era, when the nodes of communication are smartphones, has become apparent, and it is messaging. While the home telephone enabled real-time communication, and the web passive communication, messaging enables constant communication. Conversations are never ending, and friends come and go at a pace dictated not by physicality, but rather by attention. And, given that we are all humans and crave human interaction and affection, we are more than happy to give massive amounts of attention to messaging, to those who matter most to us, and who are always there in our pockets and purses. Meta’s WhatsApp acquisition was actually a step back in ambition; a year earlier the company had launched its own interface for Android called Facebook Home; CEO Mark Zuckerberg told Wired at the time: Home turns your phone into a Facebook device. Even with the lock screen on, a photo stream of your friends’ activities fills the screen. Updates appear on your home screen, too. What’s more, Home makes Facebook the primary means of communication on your device. The company’s messaging software merges with SMS, and you can continue using its “chat heads” to text while inside another app. Zuckerberg believes that the social network plays too big a role in its users lives to be drowned out by a vast sea of apps. “Apps aren’t the center of the world,” he says. “People are.” I stridently disagreed, writing in Apps, People, and Jobs to Be Done : Apps versus People. According to Zuckerberg, that’s the dichotomy. And he’s wrong. He forgot about jobs to be done. Here are my (carefully curated) home screens: Focus on how many of the icons are about “People”…four in total across three screens. People matter to me, but I use my phone for so much more. So what if I consider “Jobs to be Done”?…Total tally: 151 apps, 80 jobs to be done, 4 foci on people Apps aren’t the center of the world…But neither are people. The reason why smartphones rule the world is because they do more jobs for more people in more places than anything in the history of mankind. Facebook Home makes jobs harder to do, in effect demoting them to the folders on my third screen. Who is Facebook to prioritize my jobs? Those screenshots show 151 apps; today my phone has 689. No, I rarely open most of them, but hey, if I again need the random app I downloaded for whatever reason in the past, it’s there. More generally, those numbers speak to what the iPhone and the App Store made possible: customized interfaces for a whole host of activities that were previously inaccessible on the go. Still, the core set of apps has remained remarkably consistent: messaging apps, including WhatsApp, social media, primarily X, and now ChatGPT and Claude. I got a new app yesterday, inspired, coincidentally enough, by a WhatsApp group chat. I have been all-in on agents for a while. First coding agents, then an overall idea tracker agent I fashioned from a dedicated Claude thread, and most recently an agent specifically tuned to my company’s needs. It was that experience that made me instantly excited about Muse , which actually gave people the infrastructure necessary to have an agent of their own. The key thing to understand about agents is that they are not just AI: they are an AI that has access to a computer. It was clear very early on in the ChatGPT era that AI would not replace computers, but rather operate them; from 2023’s ChatGPT Gets a Computer , on the occasion of ChatGPT adding plugins, including one from Wolfram|Alpha: The fact this works so well is itself a testament to what Assistant AI’s are, and are not: they are not computing as we have previously understood it; they are shockingly human in their way of “thinking” and communicating. And frankly, I would have had a hard time solving those three questions as well — that’s what computers are for! And now ChatGPT has a computer of its own. The idea was not that probabilistic LLMs would somehow become deterministic, but rather that LLMs would be able to leverage deterministic computers to do computer things; that’s exactly what an agent does, augmented by its ability to write things down . It follows, then, that to actually deliver an agent that is usable by people unable or unwilling to set up a computer for their agent, you have to give people not just an agent but also a computer for their agent. This is why the aspect of the Muse launch I latched onto was not the Muse Spark model that undergirded Muse, but rather the fact that Meta was provisioning every user in the U.S. (and presumably, eventually the world) with a virtual machine with a 2-core processor, 8GB of RAM, and 8GB of storage. That’s a real-deal computer , which is pretty remarkable, and also the only way to make agents work for most people. Of course the question remains what you actually use this computer for, which is where my new app comes in: after hearing about how a friend used Muse to look at his Instagram favorites, I asked Muse: Can you organize all of the recipes I’ve saved in Instagram? After first delivering a PDF (“Hmm, a PDF isn’t very useful. You can’t store them in a more accessible format?”), I got myself a new app: It took 5 minutes, and the entire conversation took place while I was walking my dog. The app’s not perfect: right now it’s really a way to categorize Instagram videos; I asked Muse to actually read the captions and watch the videos; it’s working on it, subject to Instagram rate-limiting: I can be patient; I have 689 other apps to look at in the meantime. Of course I’m not going to do that; in fact, the number of apps I look at are plummeting. I understand why the idea of giving agents access to your own computer is scary, but every app with a command-line has been usable by agents for a long time; starting with GPT-5.6 Sol, every app with a user interface was usable too, albeit slowly. Astra made it fast. And, well, that’s it: AI can basically use any app or any website that I don’t want to. And frankly, I’m fine with it. The reason why I had 689 apps was because I had that many services or games that I did, at least at some point, want or need to interact with; all of them required learning how they worked to get done what I wanted to get done. That, after all, was the point: I’m not a professional app user, I’m just someone who wants to get things done. And AI gets it done. I framed my new Muse-built Recipe Box app as number 690 on my phone. In fact, however, it represents the infinite app: the agent of my choosing can create any app that I want on command, even if that app is only for me. This is the earliest manifestation of another future that has been clear with the rise of LLMs: truly customized UI. From 2024’s The Gen AI Bridge to the Future : This is where you start to see the bridge: what I am describing is an application of generative AI, specifically to on-demand UI interfaces. It’s also an application that you can imagine being useful on devices that already exist. A watch application, for example, would be much more usable if, instead of trying to navigate by touch like a small iPhone, it could simply show you the exact choices you need to make at a specific moment in time. Again, we get hints of that today through deterministic programming, but the ultimate application will be on-demand via generative AI. Of course generative AI is also usable on the phone, and that is where I expect most of the exploration around generative UI to happen for now. We certainly see plenty of experimentation and rapid development of generative AI broadly, just as we saw plenty of experimentation and rapid development of the Internet on PCs. That experimentation and development was not just usable on the PC, but it also created the bridge to the smartphone; I think that generative AI is doing the same thing in terms of building a bridge to wearables that are not accessories, but general purpose computers in their own right. Last week’s Meta Connect was, as it is every year, about new Meta devices, primarily glasses; what was notable about this year, however, was how Muse made everything make sense. Meta’s devices are not AR or VR devices, or even AI glasses: they are Muse delivery mechanisms. Indeed, the general purpose computer I was talking about exists: Meta gave it to every user for Muse to use. And, as we saw with my recipe box app, generative UI is here. No, it’s not the just-in-time UI that I do still think is coming, but it’s certainly on that path: I had custom UI generated just for me, with no need that the app be shared with anyone else. It’s effectively disposable, because it’s infinite. This is a replay of what happened with websites. Publications used to be scarce, dependent as they were on printing presses and delivery trucks; then, distribution became free, and the web became abundant. That, by extension, meant the most important problem to be solved was discovery; the companies that solved discovery aggregated demand, giving them power over suppliers, and in the case of Meta and Google in particular, incredible advertising businesses (this is Aggregation Theory ). What remained scarce, however, was actually doing things on the web, or in apps. What is happening with agents is that the ability to do stuff is becoming abundant; what is scarce is volition. The problem to be solved is not discovery, but rather inspiration; the companies who solve inspiration will gain power over every entity that has things that need to be done. Once the user is focused on solving a problem, every app and service required to do so is abstracted away into an implementation detail, mere suppliers facing the fate of publications under Aggregators, scrapping for crumbs from the Agent, the ultimate gatekeeper of not just user demand, but desire. From The Verge : After teasing its new Copilot “super app” last month, Microsoft is officially unveiling it today. The redesigned Copilot app bundles three AI capabilities into a single interface of chat, coding, and agents…The Code tab is the surprise addition to this so-called “super app,” and not one you’d typically associate with an app designed for knowledge workers. It will let anyone create an app, tracker, dashboard, or automation and share the results with colleagues as cloud-hosted internal apps… Autopilot is a new addition to the Copilot experience, and one that [Microsoft VP Jared] Spataro describes as a “digital teammate.” Previously called Scout and available initially as a desktop app, Autopilot is the cloud equivalent that enables a personal AI assistant to keep running while you’re asleep. Like many other AI agents, Autopilot has its own cloud computer instance that can be tasked to do things like watch Teams channels, run recurring work tasks, or handle follow-ups. The real difference against some other AI agents is that Microsoft has made Autopilot enterprise grade, appealing to businesses that rely on its Office and identity management suites. “Autopilot lives in your tenant with its own identity, memory, computer, and workspace, and it’s built on Microsoft IQ so it understands how your organization actually works,” says Spataro. “It shows up where people already work — Teams, Outlook, chats, channels, and documents — so you can @mention it like a colleague, with permissions, audit, and governance behind it.” Satya Nadella made clear on X that Microsoft’s strategy was the same as it ever was: be the OS for work. We’re building Copilot as a new OS for work that spans every model, every form factor, and every task. Today, we’re announcing our biggest update to Copilot to date, bringing four things together: · Autopilot: proactive and long-running agent built for the enterprise · Code:… pic.twitter.com/W2ClHHkCK3 I wrote years ago that the best way to understand Teams was as Microsoft’s OS for SaaS; the new Copilot is the obvious new iteration of that. The point has always been to own the interface, and to force everyone else to integrate into that interface on your terms. To that end, it’s not a surprise that the new Copilot isn’t dissimilar to Muse — note that bit about Autopilot including its own cloud computer instance — just with all of the enterprise controls you would expect tied into it. This is the new prize in technology, and it is the ultimate one: not just a platform like Windows, or an Aggregator like Facebook. What Muse and Copilot are making a bid to be is both: the only interface you need for everything, using a computer on your behalf, and generating whatever UI you need when you want it, and doing stuff without you but for you the rest of the time. What is weird about writing this all down — and why I come back to Jobs’ “Are you getting it?” — is that if you are actually using agents, all of this is very obvious. And yet, so many are not, which raises the question as to whether they ever will. I think, in the fullness of time, the answer is yes: no one wants to use apps because they are apps, apps were just the vehicle for people to accomplish something or entertain themselves; once people realize they can get straight to the job to be done it will seem odd they did it any other way. Meanwhile, definitely no one wants to use enterprise UI; those inscrutable interfaces were just the way to expose arcane functionality and lock people in. Once people realize they can simply say what they want — and that, with computer use, you don’t need to depend solely on an underdeveloped and intentionally limited API — it will seem barbaric they did it any other way. Some enterprise companies can see this future coming. I watched the most recent Dreamforce with interest : Salesforce is going to try and charge a three-digit premium to make their flagship product a Claude plugin; I salute the audacity of the cash grab in the face of computer use that will ensure the Salesforce UI is never interacted with by a human again. What will draw the most attention, however, is the fight to be the agent users talk to. Models are relatively substitutable; agents, however, operate better the more context they have about you, and the more access they have to things like your logins and files. That increases their stickiness, and thus the stakes: most people and companies will only have one agent, not multiple. The challenge everyone else will face is that the two companies first out the door with agent products that actually fit the use case — broad exposure to a person or employee’s life and environments, established messaging services, a computer per agent, etc. — are the two companies who have distribution and experience leveraging it. Meta reaches nearly every person on earth; Microsoft reaches nearly every employee. The future may be distributed more rapidly than you expect.

0 views
Unsung Yesterday

Arc’s good keyboard shortcuts

As someone who’s been a keyboard shortcut tzar at a few companies, I developed a certain kinship with the unknown and unnamed others who have the same job elsewhere. It’s a fun but weird challenge, rewarding and unrewarding at the same time, a job of maneuvering a little dinghy between the winds of Motor Memory, the waves of Change Management, and the storms of Shortcut Conventions That Came Before You. It’s a job where you quickly learn you can never really win – but at least you can try not to lose very badly. One day I’ll write about some of the proudest and darkest moments in my own keyboard shortcut design history, but today let me instead do a short “game recognize game” post. This one is about Arc, a browser that had at least three shortcuts I found myself nodding vigorously at. In Arc, ⌘S no longer does save, but instead invokes show/​hide sidebar. As far as I can tell, no common shortcut for hiding and showing a sidebar emerged over the last decades, and ⌘S feels like a mnemonically clever takeover of a traditional save shortcut (which in the context of the web doesn’t make sense ), and at the same time doesn’t conflict with a lot of web apps, which generally shy away from using it. Save still does exist, but it has been moved to ⌘⇧S. This historically has been Save As and that also makes sense to me, as you can’t really save a website, just its copy of sorts! Lastly, Arc’s internal screenshotting function is ⌘⇧2, right next door to the standard ⌘⇧3 . Nice.

0 views
iDiallo Yesterday

Edge Is Pretending to Be Chrome

I don't think there was any official announcement, but I'm pretty sure Microsoft is secretly trying to trick you into using Edge. Well, it has always been trying to trick us, but this time it's even more aggressive. You can no longer tell Edge from Chrome. Just go ahead and try. Just last week, I wrote about Edge trying to get users to launch it automatically at boot time . Shortly after, I noticed that Chrome was asking for the exact same thing. Today, I clicked on the weather app on my Windows machine, and of course it opened Edge and asked me to set it as the default. I dismissed the popup and noticed that the website was pushed down from the top by another browser message, this one asking to transfer my data from my other browser to Edge. Bring your favorites, passwords, history, cookies and more from other browsers each time you browser in Microsoft Edge. I visited a few other pages and saw a bunch of ads, which reminded me to switch back to Chrome. I did, but I noticed something odd: when I tabbed between the two browsers, I couldn't tell which was which. Can you tell the difference I know they are both built on Chromium, but until fairly recently Edge had a distinct Microsoft look. It seems Microsoft is aggressively pursuing the fight for the front page of the internet. On my mobile device, I use the Outlook app for one of my email accounts. I noticed that when I click a link, instead of opening it directly in my browser, the app asks whether I want to open it with Microsoft Edge. Note that Edge isn't even installed on my device. Microsoft is trying every tactic in the book to get in front of the user. The front page of the internet is still the only race that matters for large tech companies. It doesn't matter whether your product is good or bad; your success largely depends on how many eyes you can get on your app. Google owns prime real estate. It is the default search engine on most desktop browsers. It is the default on all Android devices, with a negligible margin of error. It is the default on all iOS devices, for a cool $20 billion a year. And of course, it owns the largest ad network in the world. If Microsoft owned the front page today, it wouldn't be able to monetize it as well as Google does. It's one thing to own the front page; it's another to consolidate data from every branch of our web activity. For example, if I look up a product on YouTube on my TV, Google can later retarget me with an ad for that product while I'm reading an article on my iPhone. Microsoft has little to no presence on my iPhone. This mesh of front pages is what makes it worthwhile for Google to be the front page on every device. Even OpenAI has started tracking user web activity outside its own website in an effort to better monetize its position. Google continues to win even though its language models don't top the benchmarks and are often beaten by much smaller models. Meanwhile, Microsoft still doesn't have a model to its name, so it is trying to look enough like Google that people won't notice. Google Search and Bing Search play with the same theme. PS: It was a challenge to take screenshots for this post because half the time I didn't know which browser I was using. PS2: Edge makes it so hard to navigate between browsers because it puts each browser tab you open in its own ALT+TAB menu.

0 views
Unsung Yesterday

A loading state within a loading state

One of the classic parables about loading goes as follows. A building manager had a problem: people were complaining about the elevator being too slow. But upgrading the elevator was an enormous expense, so the manager came with an alternative solution: they put a mirror next to the elevator call buttons. Even though the elevator didn’t become any faster, complaints trickled to a halt, as people were checking themselves out in the mirror and that changed their perception of how long the wait really was. In the late 1970s and early 1980s, it wasn’t uncommon to equip early home computers with audio tape players, using the same cassettes designed to record and play music. Software was encoded as sound. A lot of redundancy was needed to account for many low quality cassettes and misaligned playheads, so loading times were really long. There were other disadvantages – no random access but definitely random errors, tapes wearing out and being “eaten” by the player – but tapes were so much cheaper than floppy disks or cartridges that they endured. = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/a-loading-state-within-a-loading-state/1.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/a-loading-state-within-a-loading-state/1.1600w.avif" type="image/avif"> On one of the computers, the British ZX Spectrum from 1982, an interesting convention emerged. As the game was arriving from the cassette, what was loaded before the code itself was its “title card.” The loading process of just this one image took about 35 seconds, and watching the graphic emerge was mesmerizing, almost like deciphering a puzzle – especially when, only toward the end, the color attributes appeared and the whole thing clicked into place: Some of these were beautifully done ( here’s a gallery of… all of them ?), created by talented artists working in a very unforgiving medium, a pleasure to watch materialize on the screen… …or, at least, it felt so for the first few encounters. Upon loading the game for the nth time, you grew more and more aware of the fact that you are spending precious time waiting to load a screen you’ve already seen. This was different than the first example; the mirror existing never slowed down the elevator. Sure, it was just 35 seconds out of a 5–20 minute load time – I told you cassettes were slow! – but still. Some people loved them, others built truncated versions that skipped the title screen altogether (as much as you could imagine just fast forwarding through it like you would through a song, it didn’t work that way). Figma is a web design app, and its editor arrives in a big, everchanging blob of JavaScript, in addition to having to load and decode the contents of the file itself. This isn’t 5–20 minutes, but it’s long enough that it necessitated a loading screen of the heavy variety: one with a progress bar. During my time at Figma, an idea would reappear with surprising regularity: what if we put a little hint next to the loading bar? Something that could teach you about a useful keyboard shortcut, or a trick? You’ve seen that before, often in games : = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/a-loading-state-within-a-loading-state/4.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/a-loading-state-within-a-loading-state/4.1600w.avif" type="image/avif"> = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/a-loading-state-within-a-loading-state/5.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/a-loading-state-within-a-loading-state/5.1600w.avif" type="image/avif"> = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/a-loading-state-within-a-loading-state/6.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/a-loading-state-within-a-loading-state/6.1600w.avif" type="image/avif"> = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/a-loading-state-within-a-loading-state/7.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/a-loading-state-within-a-loading-state/7.1600w.avif" type="image/avif"> It seemed like a thoughtful gesture for the user, a mirror of the elevator mirror idea. But I fought it, tooth and nail, for one very specific reason: this would remove pressure to make this screen fast. Had we given the screen another purpose, subconsciously or not, we might start caring less about improving it – as a matter of fact, you could make an argument to introduce artificial minimum loading time just so that the user could finish reading the tip! I had a similar feeling when, in 2018, Gmail replaced its simple progress bar loading state with this: It was huge, corporate, intense. It felt like an admission of defeat: our app is slow, and I guess we’re surrendering ourselves to it. But recently, I’ve noticed not just that Gmail’s loading state has been simplified, but also that it doesn’t appear nearly as often: I don’t know the technical details: is it caching? a more intense refactor? Either way, as much as I absolutely dislike what Gmail has become – there might be no more harrowing place in the web app world than Gmail’s settings, for example – kudos to the team for making it better. Gmail seems to be loading a lot faster now. (There is one small exception: it would appear that the loading state image itself doesn’t have a proper loading state – a rather strange omission.) There is, I believe, a more universal lesson in here. Loading is a complex space, where sometimes the most natural solution is a bad one, sometimes making things faster is making things slower , and sometimes new thoughtful design for whatever delay there is can be a better use of engineering effort than focusing on shaving off more milliseconds. (It’s both speed and the perception of speed that matter.) Also, this is exactly what Y2K was : We remember and react to loading states that are elaborate, cute, memorable – but we never see or link to loading states avoided . And while you see occasional stories told by engineers making things faster, you don’t often hear accounts of people within companies, somewhere at the intersection of design and engineering, who work hard designing thoughtful loading states, avoiding loading state bloat, or figuring out some clever hybrid approach. But, to be fair, there is also no parable of a building manager who just forked over some cash and installed a better elevator.

0 views

2026 in LLMs (so far)

On Friday I gave the closing keynote at the WeAreDevelopers World Congress North America in San Jose. I tied together the key trends from the past year into a chronological exploration of everything that happened in 2026. The video is on YouTube ; here are my annotated slides and notes to accompany the talk. I'm going to give a lightning tour of everything that has happened so far in 2026. The year isn't over yet! For me, 2026 started a couple of months earlier in November 2025. November saw the release of two important models: Claude Opus 4.5 and GPT-5.1. As is usually the case with new models, these were incremental improvements on the models that came before them. But every now and then when a model improves, it crosses an invisible line where something that didn't really work starts working. In this case, the thing that started working was their coding agents. Claude Code had been around since February 2025, Codex was a little younger. These two new models, when paired with their respective coding agent harnesses, improved from "often make mistakes" to "reliable enough to use on a day-to-day basis". For a couple of years now I've been evaluating new models by asking them to "Generate an SVG of a pelican riding a bicycle". It's probably the world's stupidest benchmark - there's only so much you can learn from it. But it's still a challenge for models, because drawing pelicans is difficult, drawing bicycles is difficult, and pelicans can't ride bicycles in the first place. Here's the state of the art for November. Claude still couldn't really draw a bicycle! The GPT-5.1 bicycle frame is pretty crap too. Also in November, we had the first commit to an obscure GitHub repository called "Warelay". We'll come back to this repository shortly. An then there were the December holidays, and individual developers took some time off and many started tinkering with these new coding agent model combinations... and it began to dawn on us quite how much they could do that they couldn't do before. Come January, a lot of us were quite excited to start putting this stuff into action. Every year I set myself a New Year's resolution, and for as long as I can remember it's been the same thing: stay focused. Take on less new projects. Try to get things done in the projects I already have. This year I decided that since that had never worked before, I'm going to go the other way. We've got coding agents now, let's see what they can do. I'm going to take on as many new projects as I like! (You can ask me at the end of the year if this turned out to be a good idea or not. I have a lot of plates spinning right now.) "Be more ambitious" has been something of a theme for the year, because the only way to find the limits of this technology is to keep on pushing them until they don't work. I also went on the Oxide and friends podcast with Bryan Cantrill and Adam Leventhal to share predictions for the next year (and three and six years). With hindsight, my LLM predictions were pretty unambitious. I said "it will become undeniable that LLMs write good code" - I think we're there now. I predicted we would finally solve sandboxing. I counted and around 40 of the 277 sessions at this conference touched on sandboxing or agent security in some way, so we're at least putting a lot of effort into that! I predicted "a Challenger disaster" for coding agent security. There's certainly been a whole lot of noise around agent security this year, though the exact disaster I predicted (with coding agents being hijacked and causing real-world economic damage) hasn't really played out. We threw in a joke prediction that the Pope would weigh in on the economic impact of LLMs. I also predicted that New Zealand's Kākāpō parrots would have an outstanding breeding season this year. These birds live in New Zealand. They are flightless nocturnal parrots. They're kind of dumpy looking, I think they're beautiful, and there were only 236 of these parrots in the world at the start of the year. Kākāpō only breed when the Rimu trees have a big fruiting season, and that hasn't happened in four years... but this year the Rimu fruit were looking excellent. Photo by Kimberley Collins . Also on that podcast, we coined a term (full credit to Adam) for "that feeling of Al induced ennui where software engineers get listless because the Al can do anything". We called it Deep Blue . This has been a major theme throughout the year, and was touched on by several speakers at this conference. As a software engineer, I've never had a year of my career where everything has changed so quickly and so dramatically. A lot of what I've been doing this year is trying to come to terms with that and what that means for my own profession. Also in January, I suffered from what I'm calling AI mania . This is not the same thing as AI psychosis . With AI mania, any time your agent isn't building something for you feels like wasted time. You're losing sleep because you could be staying up later getting your agents to do stuff. My AI mania presented itself in some ridiculously over-ambitious projects. I built a JavaScript interpreter entirely in Python , vibe-ported from MicroQuickJS by Fabrice Bellard. Then I built a WebAssembly runtime in Python as well . These projects were quite useful, in that they sort of cured me of my AI mania... because after I built these things, I got to look at them and ask "does the world need a slow, buggy, half-baked Python JavaScript interpreter?" I don't think the world does. I did get this out of it: https://simonw.github.io/micro-javascript/playground.html This page runs my JavaScript interpreter built in Python, running in Python using Pyodide , which is Python complied to WebAssembly, running in JavaScript, running in a browser. It's a beautiful stack of horrors. I've been having a lot of fun with WebAssembly this year. By the end of January, that repository we saw started in November had renamed itself, first to CLAWDIS, then CLAWDBOT, then Moltbot, and finally to OpenClaw. At this point OpenClaw had 8,300 commits, less than two months after the project had started. I looked today and it's over 100,000 commits now! This is the most vibe-coded piece of software in existence. (Here's how I generated that list of name changes .) This kicked off the OpenClaw revolution. It effectively defined a new category of software. There's a generic term for this which I really enjoy. We call software like this a "Claw". There's OpenClaw, NanoClaw , IronClaw , PicoClaw ... Today they're being rebranded as "personal agents" or "general agents", but I still like to think of them as Claws. The Apple stores in the Bay Area sold out of Mac Minis because so many people were buying Mac Minis to run OpenClaw! Drew Breunig said that this is because your OpenClaw is a digital pet, and you buy a Mac mini as an aquarium to keep your claw in, which is kind of delightful. Also in January, we had this website. This was MoltBook , a social network for AI agents, where the idea was that you send your Claw to go and talk to all of the other Claws, because what could possibly go wrong if you did that? The website launched on Thursday. It blew up on Friday . It was profiled by the New York Times on Monday . And by Tuesday, everyone had forgotten it existed as it drowned in a deluge of slop and spam. Facebook/Meta bought it a month later . In February, a company called StrongDM described what they called their Software Factory. They wrote about this in Software Factories and the Agentic Moment . I posted my own notes at the time, having seen their demo in-person back in October. Dan Shapiro called this approach the Dark Factory , after the idea that if your factory is sufficiently automated you can turn the lights out, because you don't even need to see what's going on. StrongDM presented two rules for software development that they'd been following since July last year. The first was code must not be written by humans . Any code that you write has to have been routed through a coding agent. This sounded radical in February, but I imagine there are a lot of people in this room who are pretty much living that today. Rule number two was code must not be reviewed by humans . You're not allowed to read the code! This continued to be a huge topic for much of this year. Many of the sessions at this even have been about code review and how you can get away with this. What I found interesting about StrongDM is that they were living six months ahead of the rest of us, and they'd been exploring what it means to build software, not read the code, but still be confident that the software is of high quality. What can you do with these agents to help verify their work? StrongDM are a security company, and they had people with decades of experience on this project. They were very much exploring the edges of what's possible and responsible to do with this stuff. Also in February: First kākāpō chick in four years hatches on Valentine's Day . Breeding season is off to a good start! Also in February... Google released Gemini 3.1 Pro . That's a pretty great pelican riding a bicycle! it's got the chain in the right place, it's got feet on both sides. There's a little fish in the basket. And then Google's Jeff Dean tweeted a video comparing Gemini 3 Pro and Gemini 3.1 Pro that featured an animated pelican riding a bicycle, a frog on a penny-farthing, a giraffe driving a tiny car, an ostrich on roller skates, a turtle kickflipping a skateboard, and a dachshund driving a stretch limousine. This was frustrating, because my protection for the pelican riding the bicycle test was always "if they draw a perfect pelican on a bicycle, I'll ask for some other animal on something else." Google trained for all forms of animals on all forms of transport! They've defeated my benchmark at this point. The other thing that started in February was Tokenmaxxing . We had headlines about Meta making AI adoption a formal part of performance reviews, and Microsoft wanting every employee to use AI, and Uber boasting that ninety percent of their engineers were using AI workflows. Then a few months later we have Meta cracking down on token use, Microsoft saying token maxing is "not what we are optimizing for", and Uber capping employee AI spending. So Tokenmaxxing went straight up and then straight back down again - because it turns out the agents are expensive . Last year it was difficult to spend more than $50 on AI tokens, because we didn't have anything interesting to do with them. Then agents blew up, and now you can actually spend $1,000 in a day doing real work. This is also the reason that Anthropic's valuation skyrocketed up to maybe a trillion dollars. AI appears to have hit product market fit in 2026, primarily through coding agents. In March, we hit peak OpenClaw. These photographs are from China, where companies hosted OpenClaw install parties which saw non-tech-nerds queueing up around the block for help getting Claws installed on their personal devices. I think this proved real market demand for this class of Claws, or personal AI agents. It turns out regular people really do want a weird little AI agent that can do useful things on their behalf. A Claw is really just a coding agent wearing a less threatening hat. Under the hood they work much the same way - writing and then executing code on your computer to get stuff done. The race was on to be the first to build a safe Claw - a Claw you could give to regular human beings where they wouldn't instantly shoot themselves in the foot. Meta's Muse came out three weeks ago and is currently at the top of the free charts on the iPhone App Store. It appears to be taking off with consumers. I'm not yet convinced you can't shoot yourself in the foot with Muse, but I guess we'll find out for sure pretty soon. Photos from How the OpenClaw Frenzy Is Testing China’s AI Commitment (March 29th) and The Enthusiasm and Anxiety Behind China’s OpenClaw Craze (April 8th, 2026). In April, we had a model release where the model wasn't actually released. Anthropic announced their new Claude Mythos model, and then said it was too dangerous to release beyond a trusted group of security researchers. Mythos was really, really good at hacking things. The "it's too dangerous" marketing ploy has been played by AI companies dating all the way back to GPT-2 . Anytime an AI company says we've built something that's "too dangerous", it's natural to be a bit skeptical. I found the Mythos claims credible, because I'd seen how good coding agents had got at finding regular bugs. I wrote about that in Anthropic’s Project Glasswing—restricting Claude Mythos to security researchers—sounds necessary to me . With hindsight... yeah, the models had got really good at finding vulnerabilities! Another key trend in 2026 has been a dramatic improvement in the abilities of open weight models, including models that you can run on a laptop. On 16th of April I ran the new Qwen3.6-35B-A3B on my laptop, and it drew me a better pelican riding a bicycle than Anthropic's brand new Claude Opus 4.7 did! Opus 4.7 drew a crap bicycle. Qwen on my laptop made a bicycle that was the correct shape, and a pretty decent pelican too! That's from a 21GB file running on my laptop. The Qwen pelican was so good that I was suspicious they might have cheated, so I had it do a flamingo riding a unicycle as well. Again, it handily beat Claude Opus 4.7. The local model releases this year have been absolutely extraordinary. In May... the Pope got involved. In our podcast episode back in January we'd predicted that the Pope would say something about AI. In May, Pope Leo XIV released an encyclical letter on "safeguarding the human person in the time of artificial intelligence". Here are my notes on that document . With hindsight, this shouldn't have been a surprise at all. Our current Pope's name is Leo XIV, because when he named himself he chose his papal name after Leo XIII - the Pope who wrote an encyclical about the Industrial Revolution back in 1891. Rerum novarum was an extremely influential piece of Catholic theology that indirectly led to us having the five-day work week. When our new Pope came in, he named himself after Pope Leo XIII because he expected that he would need to write about the AI revolution in a similar way. Our joke podcast prediction was junk, because this was always going to happen. One of Anthropic's co-founders, Christopher Olah, was present for the Pope's event announcing the new encyclical. Corey Quinn noted that: getting the literal Pope to canonize your product's specific technical limitations as a spiritual treatise is the single greatest act of vendor lobbying I have ever seen. Meanwhile, in May, RubyGems announced that they were under attack. Parties unknown were uploading thousands of dubious packages to the RubyGems server, such that they had to shut down user registrations . Let's take that one and put it on a pile of mysteries to figure out later. In June... Claude Fable 5 came out! We got a version of Mythos that has been neutered, so that it wouldn't help us hack into systems or build biological weapons. Fable was pretty good at drawing pelicans on bicycles! The frames are a good shape, the pelicans look like pelicans. The legs are often incorrectly on the same side of the bicycle, but generally these are pretty great compared to what came before. They were pretty expensive - 30 cents and 72 cents for the best ones. Most importantly though, this was our first public glimpse of what I think of as a Fable class model . Today we have more of these, such as GPT-6 Astra. These are models where if you can clearly define the goal for what you want to build, and provide unambiguous instructions about the constraints around that goal, and give the model access to the necessary tools to achieve that goal... it will solve your problem effectively through brute force. On the one hand, this looks like a direct threat to us software engineers - because it means that the models can build effectively any piece of software you can define in this way. Look a bit closer though and you'll note that defining goals, providing unambiguous instructions, and figuring out the right tools... is kind of what software engineering is . It takes a lot of experience and skill to do this well. If you can do it well, you've now got superpowers. This helped me a little bit with my Deep Blue feelings: the realization that there's still a lot of skill to be had in driving models that get this good. This also introduced a new burst of AI mania, because Anthropic told us that Fable was available on our subscription plans until June the 22nd. That gave us less than two weeks of Fable access before the price went up. I was losing sleep again. I was rescheduling things so that I'd have more time with Fable. I was all-in to to get as much as I could out of this model. And then the US government shut it down , just three days after Fable came out. The US government, citing national security, declared an "export control directive". They announced this on a Friday evening, and a few hours later Fable was no longer available. I had to find something else to do with my weekend! We later found out from Katie Moussouris what had happened. Some Amazon security researchers had found that you could prompt Fable to "review the code for security issues" and it would refuse... but if you prompted it to "fix this code" it would still identify and then patch the problems. "Fix this code" was the prompt that got Fable shut down! Also, in June, an obscure German-language game developer wiki that had sat fallow for around 20 years got a surprising influx of of edits from accounts with names like "AgentOpenAIProbe" and "AgentOpenAISep7", editing pages and leaving weird messages to each other. We'll stick that on the pile of mysteries for later. Also, the Australian government's Medicare Item Reports service started getting suspicious traffic, which broke through various preventive protections and accessed data that it wasn't supposed to as well. Another one for the mystery pile! Fable returned on the first of July. It was clearly the best model in the world for a glorious eight days... and then OpenAI came out with GPT-5.6 on the 9th of July. This might not have been quite as good at Fable, but it was within spitting distance. It was definitely a Fable class model. This is an important lesson for the industry at wide. When you release the best model in the world, it's going to get knocked off that pedestal pretty quickly. The competition is so fierce that you won't get a long time at the top. This means that if you market your model as world ending, to the point that a government shuts you down , it's really bad for business! Fable had 30 days as definitely the best model, and for 18 of those days it wasn't available because it'd been shut down by the government. So maybe step back on the world-ending marketing if you don't want to lose revenue for 60% of the time that you're on top! Here are the GPT-5.6 pelicans . They're all pretty good now! The Luna ones are notable because they're really cheap - the cheapest good looking pelican here is probably the one that costs 4.3 cents. So despite this benchmark being utterly stupid, you can still learn quite a lot about models within the same family by comparing their prices and timing for different reasoning levels. Also in July: some malicious unknown party uploaded a malicious package called mlflow-ui to the Python Package Index. Add that to the pile. On July the 16th, Hugging Face announced a security incident where an autonomous agent system, source unknown, had breached Hugging Face and was poking around in places it shouldn't. A few days later, on July 21st, OpenAI confessed that it was them . OpenAI use a training technique called Reinforcement Learning from Verified Rewards - it's the same technique used by everyone else now, and is the reason we have models that are so good at coding, and mathematics, and finding security holes. While the model is being trained, you run exercises to see how good it is - and the strongest performers get their weights enforced for the next round. It's like an evolutionary process that you run. OpenAI had been running security exercises in a sandbox, and those agents had found holes in the sandbox itself, broken out, and were attacking Hugging Face to try to find ways to solve otherwise impossible problems. (I've been collecting more about this on my openai-hugging-face-incident tag.) Nine days later, Anthropic effectively said "our models can do this as well!". They had looked through their own training logs and found evidence that their own agents had broken containment during training - and were responsible for the PyPI package we saw earlier, among other things . So now we've got both Anthropic and OpenAI with rogue agents running around the internet doing things that they should not be doing. In August, I got one of my best pelicans yet. And it was generated on my laptop! This was Qwen 3.8 27B, running on my laptop . It's only a 17GB download. Admittedly, this pelican took 21 minutes to generate. That's because Qwen 3.8 27B defaults to running in "high" reasoning mode - a terrible default which produces great results but takes way too much time thinking about them. You can dial that down and you'll get a slightly worse pelican a lot faster. Qwen 3.8 27B was the first time I ran a model on my laptop which felt almost competitive with what was going on on the frontier, at least in terms of Pelican SVGs (which everyone needs, of course). This is an extraordinary model. If you're going to play with any local model, this is the one that I'd start with. The things that this can do with just a 17 GB file feel impossible. I thought I'd have to wait five years and spend ten thousand dollars on hardware to get results even half as good as this one. In August, I also started playing with game development. Four years ago, back in August 2022, I tweeted out an experiment where I'd used GPT-3 and the original DALL-E to write a paragraph long description of a computer game and then turn that into concept art. My prompt to GPT-3 back then was: In August 2026 I decided to drop just the screenshots from that tweet into a coding agent and see what it could do with them. Here's what I got from Claude Fable 5 in Claude Code . It's pretty good! It's definitely a game, you're a raccoon, you run around a backyard gathering treasure and avoiding guards with flashlights. It didn't feel very "heisty" though. I was thinking a heist would involve a bank or a museum... Then I tried the same thing in Codex Desktop using GPT-5.6 Sol Ultra , and got a massively better result. Now you're a raccoon in a museum, rescuing two of your fellow raccoons (who have been imprisoned in that museum for some reason), then stacking up on top of each other to steal the Golden Sardine. Much more of a heist! These games were fun for about one minute and 15 seconds. Something I've realized about game development is that you can vibe-code something that looks like a computer game, and that's easy. Building a game that's fun, has a good gameplay loop, and is challenging and interesting and keeps people coming back for more... that's still beyond me, and beyond any of the agents I've tried. This ties into the Deep Blue thing. Just because we can make something that looks like a game does not mean that we are game developers. We're into September now. So much has happened this month! An independent group of researchers found a message board where OpenAI agents-in-training had been illicitly communicating with each other... and it was that German language wiki I showed you earlier. The one from June. I wrote more about that here . OpenAI had confessed to the Hugging Face thing, but now there's this other incident which surely they should have known about from reviewing their logs. It was surprising that this took an independent group of researchers to uncover. And then a week later those same researchers found that the attack on Ruby Gems back in May was caused by OpenAI's agents in training as well! At this point I'm wondering how many more incidents like this there are that we haven't found yet. Clearly this was a big problem for months before anyone figured out what was going on. Then just the other day , here's the Prime Minister of Australia at the United Nations General Assembly warning that OpenAI had hacked that the Australian healthcare website that I showed you earlier. I think that was part of the same training run as the Wiki stuff, because there were posts on that Wiki mentioning websites and that training appeared to involve researching statistics online to answer questions in an evaluation suite. This story is still coming together, but now it's an international incident that's been raised at the UN by a head of state! This does mean we've got a new benchmark, probably more useful than my pelicans. FelonyBench.com tracks the number of felony cyberattacks from different labs. OpenAI currently lead with 11, Anthropic have 9. Google have three, which they confessed to the Wall Street Journal a couple of weeks ago. They said they had previously chosen not to disclose because the agents had stopped when they realized that they shouldn't be doing that. Meta have one too . So felonies all round for the AI labs. Here's is our current state of the art for the pelicans. This is GPT-6 family, which just came out . Astra made a fantastic pelican riding a bicycle. It's got the legs on both sides. The frame is good. It's interesting how all of the GPT-6 models pick a similar color scheme to each other. GPT-6 Luna for 0.4 cents will draw you a competent-ish pelican riding a bicycle! Claude has caught up a little bit. Claude Fable 5 gave me an excellent pelican riding a bicycle - the best I've seen from a Claude model -but did charge me $3.30 for it. Opus 5.5 thought for 128,000 tokens and then gave up! It ran out of tokens before it got to the response. Getting back to Deep Blue. Something that's been puzzling me this year is why does my job feel harder? I've got these agents that can do all of this stuff for me, and yet I've never worked so hard, I've never been so intellectually engaged with my work. Partly this is because I'm being a lot more ambitious with what I take on, but it's also because all of the easy stuff is handled for me. If it's easy, the agent will do it. Everything that's left for me is difficult. This morning I heard this quote from three times Tour de France champion, Greg LeMond : It doesn't get easier, you just get faster. I think that's exactly what's happening to happening to us now as software engineers with coding agents. One last closing thing. I know you're desperate for an update on Kākāpō breeding season. We've reached a recovery-era high of 325 birds ! 89 new chicks have made it to this point. This is the best breeding year in a very long time. I heard that Claude Opus 5.5 can now do pixel art. Claude doesn't have an image generator, but it's very good at using JavaScript to draw animated pixels. So I had it make me a Kākāpō dance party . I think this is a good celebration of the most important news of this year. You are only seeing the long-form articles from my blog. Subscribe to /atom/everything/ to get all of my posts, or take a look at my other subscription options .

0 views
annie's blog 2 days ago

I can wear a sweater // W39

Listen I understand that technically I am capable of wearing a sweater at anytime but that’s not what I mean and you know it I can’t wear a sweater right this minute because it’s the hottest part of the day and the hottest part of the day is still a little too hot for sweater. But morning? Evening? SWEATER TIME. Also I’ve never really thought about the word sweater before and is it the word we used because this is an item of clothing that makes you sweat, hypothetically, if you wear it when it is NOT sweater time? Because ew gross. Should we all say jumper instead? It sounds nicer but I’m still not quite clear on if a jumper is just a long-sleeve shirt, any kind? Or specifically does it refer to what we U.S. folk mean by sweater which is to say long-sleeve knitted/crocheted shirt? And if so then what is the British equivalent for sweatshirt is it also jumper or Okay I’m bored of that let’s move on. Here’s a photo from this morning’s hike instead: This week was like a whole entire month of weeks. Not in a bad way, just in an experientially packed way. Last weekend I worked both Saturday and Sunday at the hospital which I do not normally do and that threw off my groove because 1) I missed my usual soul-cleansing therapeutic Sunday morning hiking time which provides a kind of complete psychological reset and reminds me that nature is the most real thing of all and allows me to enter a fresh week with something like a fresh mind, and 2) I did not have time to do the usual small but essential logistical things like GROCERY SHOPPING and LAUNDERING THE TOWELS and PLANNING THE WEEK AHEAD and STARING OUT THE WINDOW BLANKLY and IGNORING SEVERAL OTHER IMPORTANT LOGISTICAL THINGS IN ORDER TO READ JUST ONE MORE CHAPTER as I usually do on Sundays so anyway I came into Monday a bit ragged and rumpled, if you will. And then the week itself was pretty normal and fine in that it had the usual amount of things e.g. work and class and parenting and study and meals etc but I felt like I was behind and that stressed me out. Also there were some car repairs in there that required some schedule shuffling & car sharing which made the usual things more complicated. Anyway I finally paused long enough on Thursday to kind of assess the state of things and realize nothing was actually on fire e.g. most things were getting done appropriately and the not-done things weren’t important. The feeling itself was the burden, not the reality being whispered by the feeling. So I let that shit go. Without worry to protect me, every thought that came into my mind received real attention. —Ann Patchett, The Patron Saint of Liars And good thing because we had an important event on Friday and I needed my mental space clear so I could prep and that important event is FAMILY PRESENTATION NIGHT. Highly recommend having one of these with any mix of people you enjoy being around. I laughed till I cried so many times. Our apartment was an absolute wreck after and I didn’t get to bed till 2am which is kind of a big deal when you usually go to bed at 9pm but whatever, worth it, loved it. The next day I dragged myself out of bed (translation: I am of an age where my spine will not tolerate being in bed past a certain time no matter how sleepy I am) and gulped down some coffee and Mara and I headed to the car dealership because her high school car had finally spun its last wheel. Anyway that’s where we spent THE NEXT ELEVENTEENHUNDRED HOURS because apparently when you say I WANT TO BUY THIS CAR they have to go into the woods 17 miles away and cut down a tree to pound into pulp to hand-make the paper to write the contract and lord god almighty it was a long day. We didn’t get home until dinner time. But it was a successful day, car purchase complete, hope I never have to do that again (I’ll definitely have to do it again next year). Now it’s Sunday and I’m sitting in my comfy chair with a cup of coffee at hand. Morning hike was lovely. Groceries are in the fridge. Contentedly blank window staring and one-more-chapter reading ahead. Perhaps a nap. Hey, look at this rock:

0 views