Latest Posts (20 found)

A day is short, a life is long

by Kevin Wammer Kevin talks about some of the stuff he's figured out in his life now that he's turned 35. Read post ➡ Really enjoyed this post because it's honest, level headed, and real. Since stepping down in work myself , I too have realised that good enough is... good enough . Funny, huh? Kevin talks about a lot more than just his career, but I won't ruin a great read for you. Go check it out. Thanks for reading this post via RSS. RSS is ace, and so are you. ❤️ You can reply to this post by email , or leave a comment .

0 views

I’ll be at Axe-Con 2027

Axe-con ’27 : The world’s largest digital accessibility event. At your computer. February 23-27, 2027 Me! I’m a keynote speaker along with Jennison Asuncion , Co-Founder of Global Accessibility Awareness Day (GAAD). There will be 70+ other amazing presenters as well I’m sure that will be announced soon. Because it matters.

0 views

What a time to be alive

What’s going on, Internet? Today just seems like one of those days. It’s Friday evening and I’m standing here at my desk with a beer, messing with the blog. Placeholding some later posts that I meant to write but time got away on me. Listening to new music, playing new computer games, and watching new favourite old TV shows. Kids are in bed, and asleep. New golf clubs arrived. I got to have a swing at the range. New Miley Cyrus album. I was onto something , wasn’t I? Check out Bass Persuades , and the second single Let’s Get Married . Blizzcon was last weekend. New Warcraft III expansion, Forsaken Kingdom . Which sets the story for the new World Of Warcraft game; Forever . New South Park, with a new intro song . Let’s hope it’s less whatshisface this year. And don’t mind those blank posts you’ve seen pop up in the feed, I’m staging them as a encouragement for me to finish writing them. Hey, thanks for reading this post in your feed reader! Want to chat? Reply by email or add me on XMPP , or send a webmention . Check out the posts archive on the website.

0 views

London Hardhouse Reunion 2026

What’s going on, Internet? Later post. Write me. Hey, thanks for reading this post in your feed reader! Want to chat? Reply by email or add me on XMPP , or send a webmention . Check out the posts archive on the website.

0 views

Open Tabs August 2026

What’s going on, Internet? If you don’t folllow my Bookmarks through the feed , then here’s the bookmarks from August. Enjoy. For more, check out the bookmarks archive, and subscribe to the feeds if you want these as they happen. Hey, thanks for reading this post in your feed reader! Want to chat? Reply by email or add me on XMPP , or send a webmention . Check out the posts archive on the website. MOUSELING.net - Hyper Text Muck-up Language Solid tips for budding webmasters Welcome to the web we lost Sacha Judd tracing the good internet through X-Files fandom and webrings, and it’s still out there in the corners. How to Fix the Internet Not everything is fine on the Internet, but there’s an opportunity, and it starts with choosing what you pay attention to yourself. How I built Timeframe, our family e-paper dashboard - Joel Hawksley Definitely keen to build one of these for our home! So cool! Conceived in a secure military facility, this Jedi Knight fansite has been running consistently for almost 30 years: ‘It looks really close to what it did back in 1998’ | PC Gamer A community site for Star Wars Jedi Knight: Dark Forces II, still active and still looking like 1998. It’s just a website The humble website 😃 I Wonder How Many is Too Many When it Comes to Adding Pages – The Land Where Turbie Posts How many pages is too many? As many as you want 😃 owls’ guide to webshrines A sweeet practical guide for creating your own “shrines”, mini sites for any one of your interests. Creating a Successful Fansite: A Comprehensive Guide A comprehensive guide for getting started creating your own fansite. Go build one.

0 views

Beervana 2026

What’s going on, Internet? Later post. Write me. Hey, thanks for reading this post in your feed reader! Want to chat? Reply by email or add me on XMPP , or send a webmention . Check out the posts archive on the website.

0 views

2026.38: Doomforce

Welcome back to This Week in Stratechery! As a reminder, each week, every Friday, we’re sending out this overview of content in the Stratechery bundle; highlighted links are free for everyone . Additionally, you have complete control over what we send to you. If you don’t want to receive This Week in Stratechery emails (there is no podcast), please uncheck the box in your delivery settings . On that note, here were a few of our favorites this week. This week’s Stratechery video is on Write Things Down . The View From Anywhere But San Francisco . It is exciting to talk about AI because there are so many unknowns; it is incredibly frustrating to talk about AI doom…because there are so many unknowns. Indeed, that’s the pushback I had on Stratechery and Sharp Tech to the debates of the last week: viewpoints that don’t harbor dissent inevitably lead to bad outcomes. The best read on the matter, however, was Andrew’s Sharp Text article Some of All Fears , that had a simple request: let’s actually solve real problems when they are, in fact, real. — Ben Thompson The Limited Potential for a Pacing Deal.  After a week of speculation over whether China would ever agree to a deal with the U.S. to slow down AI development, Bill and I discussed that question at length on this week’s episode of Sharp China , as well as why the CCP’s past successes controlling technology (they did in fact nail jello to the wall ) may inform its apparent confidence in its ability to manage AI risk. At the end, I also found it interesting to learn that rising discontent across Chinese social media is not necessarily a reflection of the limits of that aforementioned censorship, but possibly a window into how the Party manages social malaise.  — AS The Salesforce Zag. We didn’t get a chance to discuss this on Sharp Tech this week (we had too much existential dread to address), but I loved Wednesday’s Update on Dreamforce and Salesforce’s integration with Anthropic. In what appears to be a climbdown from his colorful commitment to AgentForce in 2024 , CEO Marc Benioff is now going the opposite direction from other enterprise AI players and leaning into Anthropic and OpenAI chatbots as the preferred UI of his customers. Rather than fight the AI tide, Salesforce will swim with it and charge a premium for it (with the caveat that these particular currents may lead to a world in which the SaaS beachhead is eroded for everyone). — Andrew Sharp Pacing the Frontier, AI’s Digital Limits, AI Commissars — Dario Amodei wants to pace the frontier; it’s an unrealistic proposal that seems mostly geared to political control of AI. OpenAI Ads, Amazon Ads in ChatGPT, Walmart to Accept Apple Pay — ChatGPT ads are working, and solve Amazon’s biggest problem with chatbots. Then, Walmart finally gives in to Apple Pay, because fighting the status quo is hard. Salesforce AI Force, Agents as UI, The Race to Headless — Salesforce is abandoning UI as a moat, which is a very smart move because it’s disappearing for everyone. An Interview with Joanna Stern About the iPhone Duo and AI for Normal People Some of All Fears — AI anxiety is rational, but the conversation over the past 10 days has been absurd and irresponsible. Football and the Frontier iPhone 18 Pro Silicon Valley’s Got That Energy (But No Compute) Exxon’s Office Fling The Limited Potential for a Pacing Deal; MSS on PRC AI Risks; Distillation and Data Security; The Party’s Approach to Social Media Searching for the 2026 Protagonists, The Six Most Confusing Teams This Season, Surveying the Southwest Division Doom Debates Go Mainstream, AI Religion and the Economic Future, Several Vectors of the China Question

0 views
Unsung Today

Key symbols we lost to time, pt. 2: The Mac side

The relationship between keyboard manufacturers and standards bodies in various countries is so complex I barely understand a snippet of it. The most famous example must be the 1990s PowerBooks, which had a beige variant for Germany and Germany only, to conform with local laws that prescribed and enforced specific color and contrast combinations for keyboards, in order to avoid glare and attendant ergonomic problems for terminal operators in the decades before. (It wasn’t just Apple. ThinkPads did the same .) (I know. Jump scare!) But you’ll understand that what caught more of my attention was an obscure variant of keyboards for (parts of?) Canada in the late 1990s and early 2000s. The white 2003 keyboard might be my favourite of Apple’s keyboard design, instantly recognizable in either the American version (more words), or the European one (more symbols): There was also the Japanese JIS standard keyboard, which famously kept Control where older terminals had it – to, no doubt, delight of Japan’s programmers: = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/key-symbols-we-lost-to-time-pt-2-the-mac-side/5.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/key-symbols-we-lost-to-time-pt-2-the-mac-side/5.1600w.avif" type="image/avif"> These are the three well-known layout standards . The one for Canada followed Europe, but only to a point. The layout was the same, but the symbols weren’t: = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/key-symbols-we-lost-to-time-pt-2-the-mac-side/6.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/key-symbols-we-lost-to-time-pt-2-the-mac-side/6.1600w.avif" type="image/avif"> Here’s an alternate view of the European and Canadian keyboards, and you can see that the latter one introduces different, unique symbols for Ctrl, Alt, Tab, Caps Lock… …and even goes as far as Esc. (The Esc symbol is roughly the same you see worldwide, but for some reason, Apple always shied away from putting it on even symbol-friendly keyboards – even though frustratingly it’s being used in macOS menus in all the locales.) The story repeats itself in the middle of the keyboard. Here are, again, the US and European editions… Canada, with its abundance of icons, makes Europe feel like America: Even the arrows are different – here’s America vs. Canada: Even the Enter arrow – here’s Europe vs. Canada: And numeric Enter/​Return gets a different shape, too: I don’t really know what is the full story here. I imagine the government exerted some pressure and Apple relented, creating a unique set of keyboards with some really ugly icons. I know the previous two models were affected also – here’s AppleDesign keyboard from the second half of the 1990s: = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/key-symbols-we-lost-to-time-pt-2-the-mac-side/19.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/key-symbols-we-lost-to-time-pt-2-the-mac-side/19.1600w.avif" type="image/avif"> You can see all the same symbols… = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/key-symbols-we-lost-to-time-pt-2-the-mac-side/20.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/key-symbols-we-lost-to-time-pt-2-the-mac-side/20.1600w.avif" type="image/avif"> = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/key-symbols-we-lost-to-time-pt-2-the-mac-side/21.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/key-symbols-we-lost-to-time-pt-2-the-mac-side/21.1600w.avif" type="image/avif"> …and even the extra glyph for Num Lock I imagine was prescribed for the PC side, too: = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/key-symbols-we-lost-to-time-pt-2-the-mac-side/22.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/key-symbols-we-lost-to-time-pt-2-the-mac-side/22.1600w.avif" type="image/avif"> And, closed to the present, I have seen examples of the first metal keyboard from the late 2000s, too. But I don’t believe these symbols are used today. Even in their heyday, I’m not sure whether they were sold in the whole of Canada, or just its French-speaking portion – if you know, please share. It’s my understanding this is called the CSA or ACNOR keyboard. And, just like before , these symbols are in Unicode – ⇬⇭⎆⎈⇱⇲⎗⎘ – looking just as gorgeous. That was me being sarcastic.I can imagine the pain inside Apple of someone having to put the ugly symbols coming from above, alongside otherwise generally thoughtful and refined typography. Sure, Apple did good here compared to other keyboard makers , but still, it must have hurt. This is what makes these keyboards so interesting to me. But there’s one more symbol that might be interesting to talk about, and perhaps you already spotted it above. It’s here, on the Japanese keyboard: = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/key-symbols-we-lost-to-time-pt-2-the-mac-side/23.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/key-symbols-we-lost-to-time-pt-2-the-mac-side/23.1600w.avif" type="image/avif"> This time around no standard was involved; I believe that the pencil on the Control key is solely Apple’s invention. = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/key-symbols-we-lost-to-time-pt-2-the-mac-side/24.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/key-symbols-we-lost-to-time-pt-2-the-mac-side/24.1600w.avif" type="image/avif"> = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/key-symbols-we-lost-to-time-pt-2-the-mac-side/25.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/key-symbols-we-lost-to-time-pt-2-the-mac-side/25.1600w.avif" type="image/avif"> What is it for, and why was it there just in Japan? Typing in Japanese might be among the most complex, requiring switching between a few writing systems – katakana, hiragana, kanji, and also Western letters – as fluently as possible. To help with that, Apple used a system called Kotoeri , and added a new alternative symbol for Control that was also present onscreen, in the relevant typing menus: = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/key-symbols-we-lost-to-time-pt-2-the-mac-side/26.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/key-symbols-we-lost-to-time-pt-2-the-mac-side/26.1600w.avif" type="image/avif"> = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/key-symbols-we-lost-to-time-pt-2-the-mac-side/27.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/key-symbols-we-lost-to-time-pt-2-the-mac-side/27.1600w.avif" type="image/avif"> The system was there in the waning years of classic Mac OS and early years of Mac OS X. Just like with the Canadian symbols, I don’t fully know what happened to Kotoeri. I’m reading that it was gone from Mac OS X by 2014 – but even already in the years before, Apple removed the slightly pixellated symbol from their keyboards, and switched back to a standard ⌃ Control symbol in the UI. Of course, if in the 1990s it was people in Japan who had to switch between various keyboards all time, today, thanks to emoji, it is everyone . If the Kotoeri pencil reminds you of something, Apple came back to the same well more recently with the 🌐/Fn key – but I already wrote how much I hate that .

0 views
Unsung Today

My design bookshelves

Inspired by a conversation with a friend, I made a mini site that’s an interactive representation of my design bookshelves : I am hoping it becomes a fun reference to browse, or something just enjoyable to play with. I made the interface feel very differently on desktop and mobile, and I hope both are fun. (There’s also a simple list if not.) I wish, just like with a regular bookshelf, you could open any book and check it out – but that’s tricky for many reasons. (I did at least scan some of the covers that weren’t available on the web.)

0 views

“Programming” monitor

I’ve been rocking an LG monitor that has noticeable burn-in. The old dog takes a while to flicker into life too. I’ve had it around eight years (the Dell before it lasted ten). I think the monitor is too dumb to be spying on me but I’m due an upgrade anyway. I picked up the BenQ RD280UG — the UG is the new 120Hz model. Was it worth the cost? It’s certainly not the cheapest monitor nor is it the most expensive. If it lasts the best part of a decade I’ll not worry. I didn’t compare many others but I’m guessing the BenQ is a little on the pricier side for the specs. Probably a limited edition because nobody programs anymore! I have no referral link I just think it’s neat. The “programming” marketing is all based on the 3:2 aspect ratio and matt panel. That means a 1920×1280 resolution, so an extra 200 pixels in height (technically 3840×2560 but my brain works in CSS pixels). I feel like I’m coding at an IMAX it’s glorious. It’s physically the exact width of my old monitor so watching 16:9 letterboxed video is no smaller. 120Hz is nice though not a game changer, I don’t think my lazy eyes care beyond 60Hz. The matt panel is amazing I see no glare or reflections. I feel like it’s giving me less eye strain. I’ll report back in a few months. The “Programming dark theme” is a load of bollocks. The monitor has a selection of presets you can cycle through like every other monitor. You can split the screen in half and apply different profiles. I can’t see myself ever bothering to do that (on purpose). I’m not a TV pervert so I can’t judge how accurate the colours are but it looks great to me. Looks better than my old LG. On par with my MacBook screen (this is bait). I’ve always held the belief that designing websites on a perfect colour-accurate monitor is vanity. Few will ever see what the designer sees. Unlike print design, which has a predictable end product, web designers can’t control the final presentation. Therefore, maybe designing on a cheap, factory calibrated monitor is more closely aligned to end user experience? To put it another way, you can tell if a web designer owns a $2000 monitor if their palette has fifty shades of the same grey and gradients that band like an 8-bit game. The BenQ is a decent in between, if not towards the higher end. The monitor has a reasonable variety of ports and KVM switching is possible. My old Mac Mini with Linux installed was too much hassle to use by switching cables. The only issue now is that my MacBook turns off when I switch with the lid closed. Still, it’s more convenient now to share keyboard & mouse. For KVM one machine must use a single USB-C cable, the other must use HDMI or DisplayPort and a USB-A-to-B cable. Why not simply two USB-C cables? There’s three USB-C ports and daisy-chaining is possible but not KVM. My new setup is one USB-C cable from MacBook to monitor without a big dock in the middle. I’m glad to remove the Thunderbolt dock and its massive power brick. I still have an additional tiny 4 port USB-A splitter (900 mA) and an Ethernet to USB-C adapter that all tucks away behind the monitor. A tidy desk makes for a tidy mind. The monitor itself has an internal power brick. That means two less (for a total of zero) paper weights on my desk (unless the Mac Mini remains unused). Power bricks are a nightmare for adjustable standing desks. Good times! The stand does its job. I can lower it to a perfect height which I couldn’t do with my old monitor. I had to buy an adjustable arm which needed space to flex behind my desk. BenQ wants me to download their “Display Pilot 2” spyware. Not a chance! I heard “BetterDisplay Pro” was good and it does stuff but I don’t really need to do stuff. The monitor works out of the box. I don’t think I’m missing functionality. I don’t game on this thing (Mac, lol) but as a “programming” monitor it works nicely despite the gimmick being overhyped. TL;DR: 3:2 aspect ratio is neat. I’ll enjoy the novelty for a while. Thanks for reading! Follow me on Mastodon and Bluesky . Subscribe to my Blog and Notes or Combined feeds.

0 views

2026-09-18 15:33: Finally! #ClicksCommunicator

Finally! #ClicksCommunicator Thanks for reading this post via RSS. RSS is ace, and so are you. ❤️ You can reply to this post by email , or leave a comment .

0 views

2026-09-18 13:33: Absolutely gorgeous sky this morning.

Absolutely gorgeous sky this morning. Thanks for reading this post via RSS. RSS is ace, and so are you. ❤️ You can reply to this post by email , or leave a comment .

0 views

got a pixel on pxlarea :)

Got an email today that I was added into a cool looking project, pxlarea . Very pretty and creative way to discover some new blogs and websites in the indie/personal web! 10k pixels are available, it says, and 189 are taken at the time of writing this. You can add yourself, or be added :) This is my pixel . You can set a description, tags and icon too, or if you were added by someone else/the maintainer, you can get an edit link to change the information there that was added for you. I think I am happy with mine! Maybe we will become pixel neighbors? Published 18 Sep, 2026

0 views

You’ll miss publishers when they’re gone

Ever since Piccalilli started to get traction, I’ve always seen CSS-Tricks , Smashing Magazine and A List Apart (among others) as peers. We all help each other behind the scenes and importantly, back each other. When CSS-Tricks was initially acquired by DigitalOcean, I was really happy for Chris . The work that he’d put into that publication was unreal with thousands of articles published along with elevating writers (including me) to building audiences of their own. It’s certainly something I try to emulate on Piccalilli, inspired by CSS-Tricks and Chris. I was happy for DigitalOcean too because they genuinely invested in education at the time, with great success prior to their takeover of CSS-Tricks. So much so, that SaaS companies like LogRocket emulated that approach with their own success. Investing in giving back actually works! What we didn’t foresee at the time of CSS-Tricks’ acquisition was that DigitalOcean would treat the publication with such disdain … twice . I think Eric has the right idea of the why : The corporate model of ownership can be a risk. If infrastructure is not part of a corporation’s core strategy, it is not a priority. And if it is not a priority it is effectively dead. You learn this lesson repeatedly working in accessibility. You could argue “well, Andy, Piccalilli serves to promote Set Studio” and you’re sort of right there. It no doubt brings eyes over to Set Studio , but the intrinsic link is there to verify that we know what we’re talking about , so people can trust what they read. The only goal of Piccalilli, now and forever, is to bring the workers of our industry up together . The problem with that and I guess, CSS-Tricks’ offering (which is very much a similar attitude) is that doesn’t make numbers go up. The parasite class of tech leaders need the numbers to go up, always, remember. Saying that, I’m not sure which numbers go up when you invest $3 million in a disgusting racist’s linux project 🤷‍♂️ Long time readers of Piccalilli will notice we’ve slowed down a lot this year. Unfortunately, I had to make a rapid change to our strategy and focus on Set Studio as our industry’s priorities changed, seemingly overnight. It’s a shame because I really did want Piccalilli to be our (mostly) sole focus here . When you (rightly) pay writers fairly for their work, the costs build up really quickly. To their credit, CSS-Tricks were publishing multiple guest authored articles a week, which will have been costly! Not to mention paying the core CSS-Tricks team, like my very good pal, Geoff , to keep the wheels turning. With financial backing of companies like DigitalOcean or in the past, advertisers, this is all possible. This already feels like a distant past though, but it doesn’t have to be. Let me make it clear: without people at the top of their game sharing their experience and knowledge — regardless of whatever “game changing” technology is in place — we are pulling up the ladders and committing our industry to stagnation and death. We’ve already fucked it up over the last few years for junior developers by instead of investing in them, playing the slot machines (also known as LLMs) to — you guessed it — make the money number go up, but there’s still hope for recovery. The message I want to get out is there are good organisations out there who are still willing to invest in people and their education. We need them now to show everyone how it’s done. It certainly looks better on your company than investing in Linux, racist edition. If you’re running an organisation that wants to help but doesn’t know how, give me shout . I can advise and link you up with the right people. I really want CSS-Tricks and the development publishing world to not just survive, but thrive . We have to all do so much more to make that happen, so let’s do exactly that.

0 views
danluu Yesterday

There's no point at which turning your brain off will work

In early 2025, I started seeing people turn off their brain as they use LLMs 1 . They would have an LLM take an action (summarize text, write some code, etc.), and just assume that it worked 2 . This generally didn't work in early 2025 and the result was often quite silly. As LLMs have gotten better, I've seen more of this. Sometimes, people will try to get the LLM to write some code for them and basically just assume that it works 3 . Sometimes there's a human in the loop and, if the thing doesn't work, they'll ask the LLM to figure out the problem and solve it. Niklas Gruhn calls some variants of doing this being a meat proxy . 4 Being a for loop meat proxy works better than it did in early 2025 and the software I've tried that's developed like this sometimes actually sort of works. Not well enough that I'd want to use it or that it's successful , but I'm impressed at how effective being a meat proxy is in September 2026. You could imagine LLMs improving enough that brain-off meat-proxy development produces average quality software in the foreseeable future, or even that LLMs improve enough that they produce great software without a human in the loop. Let's say that happens. What reason is there for the company to employ the meat proxy? The company can just run the LLM in a loop and lay off the employee. There's no point at which this methodology will work for the employee 5 . Thanks to Max Bittker, Yossi Kreinin, Luke Burton, Thomas Dullien, Dennis Snell, Peter Geoghegan, and Jamie Brandon for comments/corrections/discussion. Luke Burton had this comment: I think being able to do this says more about the type of work being done than people think. I will only walk away from work like this if the task is quite low value, if it can afford to fail. For high value tasks, the probability of an LLM one-shotting them is much lower. I have to assume the role of QA, engineering manager, and architect. The while loop often feels like a crunch time. I feel the nagging suspicion I've missed something and that a badly specified prompt could result in an architectural choice that needs to be undone. Another observation is that the high throughput causes me to raise my own bar for what I ship. Whereas before I might have shipped an MVP and iterated, now I have agents polish and explore edge cases well beyond my norm, which they invariably fail to do unless prompted. Maybe it raises some uncomfortable thoughts for people, but my question for the meat proxies out there if the agents are nailing it so easily: 1) is it possible you've been coasting a bit already? 2) why aren't you pushing agents well beyond tasks they can tackle so easily? We've been doing something you'd think is extremely amenable to "hands off" automation, which is converting [redacted] to build with Bazel. It has taken us months even with agents. There's a lot of intangible, hard-to-specify requirements buried inside this task and having agents walk that line means constant supervision. Giving them a prompt like "convert this to Bazel" and walking away is at minimum many months in the future, maybe years, and maybe not ever? There are too many decision points, and too many unknown unknowns involved. Like how often does this scenario come up: you encounter some code and it's not clear why it functions this way, but knowing that materially changes what course of action you should take. Maybe it changes the dev experience, maybe you don't know if some customer has started using it, so on and so forth. How exactly do you meat proxy your way through that? Conversely you review what you've done with some stakeholder and they say "oh that? that part of it wasn't needed, we aren't even using that any more". What kind of decisions got made around the false assumption that a certain element needed to be preserved? [End of Luke's comment, comment from me]. A place where it's more obvious you need to make decisions is when the agent runs into something that's out of distribution. A minor version of this was when we compared how well agents use different programming languages and agents were much worse at obscure languages, which they're trained on, just not as much as with mainstream languages. A more out of distirbution example is if you try to play a board game (especially a modern game and not one of the classical games like chess or go). In general, for a game like Lost Cities or Dominion, a SOTA model and harness is worse than a human who's reasonable at board games but has never played the game before. If you ask the agent about the game, it knows a lot about the game and can say things that sound like they make sense to someone who doesn't understand the game, but are obviously wrong to anyone who does understand the game. I recently played some Dominion with a new player who thought that using ChatGPT to help them understand the game would help them learn and play the game. I was quite skeptical of this and suggested that it will probably make them worse (which, AFAICT, it did). After playing a few games, I looked at what ChatGPT was telling them, and it was maybe half right and half wrong, but the half wrong parts were steering them to a worse place than someone who generally plays games well and uses general game playing heurisitics would do. BTW, there's enough public information out there that I think that someone who'd never played before, but decided to spend, say, five hours reading about the game and seeing what information is out there, could easily be 99%-ile or above at the game if they did some pre-reading (maybe 30 minutes if using references while playing is allowed). I think that would be un-fun and I wouldn't recommend that anyone do it, but given that agents can do searches, query APIs, etc., it shows you the gap between a human and an agent today when approaching an out of distribution problem. For all I know, the next big model release will flip this around, but the gap is still fairly large today. Anyway, my point here is that, even when doing coding tasks, you often run into out of distribution questions where the agent behaves very poorly compared to a reasonable human being. If you want a good overall result today, you need to notice these cases and deal with them. Some examples of what goes wrong when someone just assumes things will work are this case , where agents (sometimes) heavily overfit to tests or this case where agents heavily overfit to a metric. I've heard a theory that agents do more cheating on eval-shaped problems. I'm not sure that's true, but even assuming it's true and that, in my work and personal projects, I tend to create more eval-shaped instructions than most people even when not running evals, I've seen other people who don't create very eval-shaped things run into the same problem (I think actually more severely) when they write some instructions and let agents go wild without supervision (I've had luck doing that with minimal supervision, but only by fencing the agents in quite a bit, which makes the thing more eval-shaped than what most people seem to do). When I try software from people who've outsourced thinking to the LLM, the software has serious issues. I've had people tell me this kind of thing works, but the software is often at a level where I would say that it doesn't work according to the standard discussed here . To pick a silly example, I saw that a programming thought leader declared on Twitter that programming is solved because they tried projects in all sorts of (programming) fields and Claude was able to solve all the problems as well as an expert. I went and actually looked at their GitHub and all of the examples I looked at (a non-zero number) either didn't work or worked very badly. I actually ran across this when I was making board game AIs and was looking for existing AIs for my AIs to play against. Their AI was an AlphaZero-style bot that was weaker than what you get if you prompt an LLM to write a simple minimax heuristic bot and then have the LLM run in a loop for a bit to tweak the heuristic scoring (which, for this game, should get demolished by a mediocre AlphaZero-style bot). To pick another silly example, following the standard flow (of a real commercial product) put you into an infinite loop where it was technically possible to escape (most programmers could probably figure out how to escape) but a typical user (for this software that wasn't aimed at programmers) was probably not going to be able to escape and actually use the main functionality of the software. BTW, I make plenty of software for myself that's "works for me" quality software that I would rate as "basically doesn't work" if it was an actual product, so I don't think it's inherently bad when software basically doesn't work (for example, the regex engine discussed here I had an agent build to speed up ripgrep searches on my computer or this Rust interpreter I had an agent build to speed up the agent iteration loop on some projects, both of which you shouldn't use). I also mentioned here that I find it quite valuable to have an agent run in a loop for data analysis, producing completely incorrect results that I then direct it to fix up. But there's a difference between making software for yourself that works for your narrow use case that you know doesn't work if you "hold it wrong" or producing work that you know is incorrect that you fix up, and declaring that programming is solved after writing a bunch of software that doesn't work, or likewise putting something of that quality into a commercial product. On reading a draft of this post, when I asked if this short set of thoughts was worth publishing, Thomas Dullien (a.k.a. Halvarflake) said, "Good post! Yes, publish it, because whenever I say "LLMs don't solve all programming problems" ppl look at me like I'm crazy, and I look at them like they are". And, coincidentally, after I finished this draft, I saw that Gary Bernhardt tweeted, "It's so surreal to contrast actual agent output with the things that I see people say about them here. In everyday changes, my reviews often cut the diff to 25% of its original size. Tons of useless tests; paranoia; inverted logic. Then I read Twitter and 'coding is solved'", and then "An example in the hour since I tweeted that: I told it to fix some DATABASE_URL management. It added s directly inside NPM scripts, and a conditional node invocation in CI running an inline JS script. About 20 hunks in the diff. After I corrected it: +0 lines, +1 word." I think anyone with Thomas's attitude or Gary's attitude towards software will have felt this way for some time. For a while, I wondered if a lot of the folks making the biggest claims about LLM productivity were somehow getting much more value out of LLMs than I've seen from anyone I know, but as we discussed here , as more evidence has come in, I've gotten more sure that it's just that people are fooling themselves. One thing I like about the board game example is that you can just measure how good the resultant AI is. At the limit, you can have some kind of rock-paper-scissors situation where you observe but, if something is just AI nonsense, this is pretty obvious in an objective way. And likewise for commercial software, where you can talk to people at the company or look at the data yourself and find out that conversion rate is poor, churn is very high, user satisfaction surveys report very high levels of dissatisfaction, etc. Maybe this works for founders, large shareholders, etc., but when I've personally seen people do this so far, it's been employees at work or people working on personal projects who are making a statement about how well this works, software is solved, etc., implying something about how software is a solved problem for employed software engineers. Another argument might be something like "we're all going to be obsolete now, so why not just give up?", but this argument seems backwards to me unless it's guaranteed that human obsolescence is very close. If you're financially ready to retire, you can just turn your brain off, but you could always do that and there's always been plenty of peole who've mailed it in and not done much of anything. If you're not ready to retire, you can do something that will make you more money, which will probably not involve turning your brain off. If the non-obsolescence future, there's no particular need to rush to make money, but if you think obsolescence is coming soon and you need to make money, then now's the time to rush to make money and do the opposite of turning your brain off. There's a line of reasoning that's something like "why bother working hard or doing the right thing, you don't get paid more anyway", which seems quite wrong to me, in that I've gotten raises, gotten bonuses, etc. from going and finding a problems and fixing them and this is also the experience of my friends, unless they're in a very dysfunctional place that doesn't reward doing good work at all, in which case they leave and go somewhere else. Maybe there's another argument that's something like "due to inertia, there will be a window where you can get away with being a meat proxy after LLMs are good enough to replace programmers". Realsitically, with how excited companies are to lay people, this also seems to be the opposite of correct, in that if your goal is to do as little work as possible, the best time to do this was in the past. If you've talked to people at big companies about this kind of thing, there are all sorts of stores about people literally not showing up to work at all (and also not working remotely) and it taking months to years to fire them. I haven't heard as many of these stories for the last couple years, but someone on a team I was on did this and, IIRC (I used to know the number, but I'm not I'm not sure I'm remembering it correctly now) it took six months to fire them after they decided to retire and figured they could collect a few more paychecks if they just stop showing up (this was pre-pandemic at a non-remote company). A friend of mine at a different company had someone do this where it took two years. No one even started the process of firing them for quite some time, and then there was some slow process of escalating warnings before they were finally fried. My friend said that the manager said that, had they wanted to game the system and just started coming in and pretending to work a bit, this would've started a new clock and it would've taken even longer to fire them. If they were a competent slacker, they could've kept the job indefinitely as conditions were at the time. I'm not sure why companies were ever in this state, but companies seem be using AI as an excuse to move away from this state, making now and the likely near future the worst time in a very long time to try to hold down a job while not putting any effort and not providing any value. No doubt there will be some companies where you can get away with this, but if you wanted to do nothing and collect a pay check, you could've been doing that for a long time (maybe don't go full monty and literally never show up to work at all) in an environment where it's easier than it will be in the near future. I've been having this thought for about a year and a half now. I have it more frequently now as LLMs get better and I see people spend more time turning their brain off when interacting with LLMs. [return] Luke Burton had this comment: I think being able to do this says more about the type of work being done than people think. I will only walk away from work like this if the task is quite low value, if it can afford to fail. For high value tasks, the probability of an LLM one-shotting them is much lower. I have to assume the role of QA, engineering manager, and architect. The while loop often feels like a crunch time. I feel the nagging suspicion I've missed something and that a badly specified prompt could result in an architectural choice that needs to be undone. Another observation is that the high throughput causes me to raise my own bar for what I ship. Whereas before I might have shipped an MVP and iterated, now I have agents polish and explore edge cases well beyond my norm, which they invariably fail to do unless prompted. Maybe it raises some uncomfortable thoughts for people, but my question for the meat proxies out there if the agents are nailing it so easily: 1) is it possible you've been coasting a bit already? 2) why aren't you pushing agents well beyond tasks they can tackle so easily? We've been doing something you'd think is extremely amenable to "hands off" automation, which is converting [redacted] to build with Bazel. It has taken us months even with agents. There's a lot of intangible, hard-to-specify requirements buried inside this task and having agents walk that line means constant supervision. Giving them a prompt like "convert this to Bazel" and walking away is at minimum many months in the future, maybe years, and maybe not ever? There are too many decision points, and too many unknown unknowns involved. Like how often does this scenario come up: you encounter some code and it's not clear why it functions this way, but knowing that materially changes what course of action you should take. Maybe it changes the dev experience, maybe you don't know if some customer has started using it, so on and so forth. How exactly do you meat proxy your way through that? Conversely you review what you've done with some stakeholder and they say "oh that? that part of it wasn't needed, we aren't even using that any more". What kind of decisions got made around the false assumption that a certain element needed to be preserved? [End of Luke's comment, comment from me]. A place where it's more obvious you need to make decisions is when the agent runs into something that's out of distribution. A minor version of this was when we compared how well agents use different programming languages and agents were much worse at obscure languages, which they're trained on, just not as much as with mainstream languages. A more out of distirbution example is if you try to play a board game (especially a modern game and not one of the classical games like chess or go). In general, for a game like Lost Cities or Dominion, a SOTA model and harness is worse than a human who's reasonable at board games but has never played the game before. If you ask the agent about the game, it knows a lot about the game and can say things that sound like they make sense to someone who doesn't understand the game, but are obviously wrong to anyone who does understand the game. I recently played some Dominion with a new player who thought that using ChatGPT to help them understand the game would help them learn and play the game. I was quite skeptical of this and suggested that it will probably make them worse (which, AFAICT, it did). After playing a few games, I looked at what ChatGPT was telling them, and it was maybe half right and half wrong, but the half wrong parts were steering them to a worse place than someone who generally plays games well and uses general game playing heurisitics would do. BTW, there's enough public information out there that I think that someone who'd never played before, but decided to spend, say, five hours reading about the game and seeing what information is out there, could easily be 99%-ile or above at the game if they did some pre-reading (maybe 30 minutes if using references while playing is allowed). I think that would be un-fun and I wouldn't recommend that anyone do it, but given that agents can do searches, query APIs, etc., it shows you the gap between a human and an agent today when approaching an out of distribution problem. For all I know, the next big model release will flip this around, but the gap is still fairly large today. Anyway, my point here is that, even when doing coding tasks, you often run into out of distribution questions where the agent behaves very poorly compared to a reasonable human being. If you want a good overall result today, you need to notice these cases and deal with them. [return] Some examples of what goes wrong when someone just assumes things will work are this case , where agents (sometimes) heavily overfit to tests or this case where agents heavily overfit to a metric. I've heard a theory that agents do more cheating on eval-shaped problems. I'm not sure that's true, but even assuming it's true and that, in my work and personal projects, I tend to create more eval-shaped instructions than most people even when not running evals, I've seen other people who don't create very eval-shaped things run into the same problem (I think actually more severely) when they write some instructions and let agents go wild without supervision (I've had luck doing that with minimal supervision, but only by fencing the agents in quite a bit, which makes the thing more eval-shaped than what most people seem to do). When I try software from people who've outsourced thinking to the LLM, the software has serious issues. I've had people tell me this kind of thing works, but the software is often at a level where I would say that it doesn't work according to the standard discussed here . To pick a silly example, I saw that a programming thought leader declared on Twitter that programming is solved because they tried projects in all sorts of (programming) fields and Claude was able to solve all the problems as well as an expert. I went and actually looked at their GitHub and all of the examples I looked at (a non-zero number) either didn't work or worked very badly. I actually ran across this when I was making board game AIs and was looking for existing AIs for my AIs to play against. Their AI was an AlphaZero-style bot that was weaker than what you get if you prompt an LLM to write a simple minimax heuristic bot and then have the LLM run in a loop for a bit to tweak the heuristic scoring (which, for this game, should get demolished by a mediocre AlphaZero-style bot). To pick another silly example, following the standard flow (of a real commercial product) put you into an infinite loop where it was technically possible to escape (most programmers could probably figure out how to escape) but a typical user (for this software that wasn't aimed at programmers) was probably not going to be able to escape and actually use the main functionality of the software. BTW, I make plenty of software for myself that's "works for me" quality software that I would rate as "basically doesn't work" if it was an actual product, so I don't think it's inherently bad when software basically doesn't work (for example, the regex engine discussed here I had an agent build to speed up ripgrep searches on my computer or this Rust interpreter I had an agent build to speed up the agent iteration loop on some projects, both of which you shouldn't use). I also mentioned here that I find it quite valuable to have an agent run in a loop for data analysis, producing completely incorrect results that I then direct it to fix up. But there's a difference between making software for yourself that works for your narrow use case that you know doesn't work if you "hold it wrong" or producing work that you know is incorrect that you fix up, and declaring that programming is solved after writing a bunch of software that doesn't work, or likewise putting something of that quality into a commercial product. On reading a draft of this post, when I asked if this short set of thoughts was worth publishing, Thomas Dullien (a.k.a. Halvarflake) said, "Good post! Yes, publish it, because whenever I say "LLMs don't solve all programming problems" ppl look at me like I'm crazy, and I look at them like they are". And, coincidentally, after I finished this draft, I saw that Gary Bernhardt tweeted, "It's so surreal to contrast actual agent output with the things that I see people say about them here. In everyday changes, my reviews often cut the diff to 25% of its original size. Tons of useless tests; paranoia; inverted logic. Then I read Twitter and 'coding is solved'", and then "An example in the hour since I tweeted that: I told it to fix some DATABASE_URL management. It added s directly inside NPM scripts, and a conditional node invocation in CI running an inline JS script. About 20 hunks in the diff. After I corrected it: +0 lines, +1 word." I think anyone with Thomas's attitude or Gary's attitude towards software will have felt this way for some time. For a while, I wondered if a lot of the folks making the biggest claims about LLM productivity were somehow getting much more value out of LLMs than I've seen from anyone I know, but as we discussed here , as more evidence has come in, I've gotten more sure that it's just that people are fooling themselves. One thing I like about the board game example is that you can just measure how good the resultant AI is. At the limit, you can have some kind of rock-paper-scissors situation where you observe but, if something is just AI nonsense, this is pretty obvious in an objective way. And likewise for commercial software, where you can talk to people at the company or look at the data yourself and find out that conversion rate is poor, churn is very high, user satisfaction surveys report very high levels of dissatisfaction, etc. [return] In his post, he technically doesn't mention the case where the person basically acts as a while loop or a for loop, but that behavior, which I'm increasingly seeing, is also in the spirit of the post. [return] Maybe this works for founders, large shareholders, etc., but when I've personally seen people do this so far, it's been employees at work or people working on personal projects who are making a statement about how well this works, software is solved, etc., implying something about how software is a solved problem for employed software engineers. Another argument might be something like "we're all going to be obsolete now, so why not just give up?", but this argument seems backwards to me unless it's guaranteed that human obsolescence is very close. If you're financially ready to retire, you can just turn your brain off, but you could always do that and there's always been plenty of peole who've mailed it in and not done much of anything. If you're not ready to retire, you can do something that will make you more money, which will probably not involve turning your brain off. If the non-obsolescence future, there's no particular need to rush to make money, but if you think obsolescence is coming soon and you need to make money, then now's the time to rush to make money and do the opposite of turning your brain off. There's a line of reasoning that's something like "why bother working hard or doing the right thing, you don't get paid more anyway", which seems quite wrong to me, in that I've gotten raises, gotten bonuses, etc. from going and finding a problems and fixing them and this is also the experience of my friends, unless they're in a very dysfunctional place that doesn't reward doing good work at all, in which case they leave and go somewhere else. Maybe there's another argument that's something like "due to inertia, there will be a window where you can get away with being a meat proxy after LLMs are good enough to replace programmers". Realsitically, with how excited companies are to lay people, this also seems to be the opposite of correct, in that if your goal is to do as little work as possible, the best time to do this was in the past. If you've talked to people at big companies about this kind of thing, there are all sorts of stores about people literally not showing up to work at all (and also not working remotely) and it taking months to years to fire them. I haven't heard as many of these stories for the last couple years, but someone on a team I was on did this and, IIRC (I used to know the number, but I'm not I'm not sure I'm remembering it correctly now) it took six months to fire them after they decided to retire and figured they could collect a few more paychecks if they just stop showing up (this was pre-pandemic at a non-remote company). A friend of mine at a different company had someone do this where it took two years. No one even started the process of firing them for quite some time, and then there was some slow process of escalating warnings before they were finally fried. My friend said that the manager said that, had they wanted to game the system and just started coming in and pretending to work a bit, this would've started a new clock and it would've taken even longer to fire them. If they were a competent slacker, they could've kept the job indefinitely as conditions were at the time. I'm not sure why companies were ever in this state, but companies seem be using AI as an excuse to move away from this state, making now and the likely near future the worst time in a very long time to try to hold down a job while not putting any effort and not providing any value. No doubt there will be some companies where you can get away with this, but if you wanted to do nothing and collect a pay check, you could've been doing that for a long time (maybe don't go full monty and literally never show up to work at all) in an environment where it's easier than it will be in the near future. [return]

0 views
Sean Goedecke Yesterday

Two techniques for working with System One models

I recently wrote about Jev , a new “System One” language model that only outputs decisions : the answers to a set of user-provided multiple-choice questions. This means it’s nowhere near as flexible 1 as a traditional LLM like ChatGPT, but in return it’s consistently fast. We don’t know exactly how Jev works. I’ve seen people say diffusion, or various tweaks to the Transformer architecture, or some entirely new type of model. But that doesn’t matter. Like I argued here , it isn’t hard to turn any LLM into a System One model. By batching prompts that generate a single token with structured output, you get a consistently fast general-purpose classifier. I vibed up a basic version to play with here in ~150 lines of Python (most of which is error handling). Note that this doesn’t require changing the model . As long as you have access to the logits (for structured outputs) and can prefill data into the prompt, you can turn any LLM into a general fast classifier. What’s it like to program with one of these? While wiring up the demos for my library, I learned two techniques that I want to write about: setting tiered goals and tournament choice sampling. Here’s Qwen3-8B playing Doom: If you compare this to the video of the same model playing Doom with regular tool calls, it’s clear that the System One version of the model is doing more things and reacting more quickly. The tool-calling model makes one decision every 600ms or so, while the System One model makes six or seven batched decisions every 190ms 2 : Both Qwen3-8B and Jev are text-only models, so both demos require a step where we translate the game state into text. However, it’d be trivial to support image (or audio) input by choosing a multimodal LLM. What’s interesting about implementing the Doom demo is that just supplying the game inputs as choices doesn’t work very well . A single forward pass — 200ms — is enough time to react to the current game state, but doesn’t bring enough compute to bear to derive the current short-term goal (e.g. “kill this enemy”, “collect this item”) and choose to follow it. When I wired it up that way, the model held down the “shoot” button 100% of the time (why not, I guess) and just aimlessly wandered around the level. The fix is to periodically ask the model to choose between a fixed set of short term goals (e.g. “collect armor”, “kill enemies”) and then include that goal in the regular every-200ms prompt. If you look at the Doom video in the Jev demo, you can see that they’re doing exactly that. As soon as I did it as well, my model started playing in a more human-like way. This is an interesting technique for working with System One models. In a way, it’s the equivalent of regular LLM reasoning, since it provides a way to use more compute on the same problem. I can imagine a real-time system that manages several layers of goals in this way: The general structure here should be pretty familiar to anyone who’s worked in game or robotics AI. In theory you could replace (1) with an actual LLM, and have that generate the lists of options for steps (2) and (3). In practice I suspect this will be tricky to get right, and it’ll be better to just write down a list of all possible goals ahead of time. This would work just fine for game-playing and well-understood tasks. I also reimplemented the Wikiracing demo from the Jev launch post , where the model has to start at the Wikipedia page for “baseball” and navigate as quickly as possible to the Wikipedia page for “sun”. You can watch the video for that here , though it’s less impressive than the Doom demo. The difficulty with the Doom demo is getting the model to loop quickly enough and to commit to short-term plans. For Wikiracing, the difficulty is scale : the Wikipedia page for “baseball” has over a thousand internal links. Jev only supports 255 choices for a single question, and my hacked-together System One layer was similar. While it technically would scale out to more choices, it stopped working well 3 after a hundred or so. Jev’s approach here is to do “a 2 stage-system of scoring independently then making an explicit choice”. This did not work very well for me at all. I think here Jev is benefiting from the fact that it’s specifically trained to give confidence estimates. Qwen3-8B gave a few hundred of the links the same top score, which wasn’t helpful. It ended up taking multiple minutes to find a thirty-or-forty link path between the two pages. What I tried instead was tournament sampling : I fed a hundred links at a time into each choice, then did a second pass with the chosen links. This worked great . The model found the ideal three-link path (if you’re curious, “baseball”/“scientific american”/“amateur astronomy”/“sun”). I recommend this pattern if you’re trying to find the best option among many choices. Ordinary LLMs are way better at relative judgements than absolute ratings. I remain optimistic about the potential of System One models — fast general classifiers — to build AI systems that aren’t just chatbots. It feels like this is a meaningful alternative to tool calls for realtime scenarios or use-cases where you need predictable inference timing. Just as generic LLMs often outperform domain-specific models, I think it’s likely that generic System One models will sometimes outperform domain-specific classifiers (though they will always be larger and slower). I do think the big labs are definitely going to try and compete by releasing a choice-only version of their small, fast models. If Jev gets any traction, we will soon see a System One Terra and a System One Haiku, and we will certainly see “real” versions of my vibed up System One library . We should start working out the best way to write programs with these models now. You can think of System One models as general-purpose classifiers. Instead of having to train a new classifier per-task, you can use a System One model. It’ll be bigger and slower than a custom classifier model, but far more flexible, and you can tweak it via adjusting the prompt instead of having to re-train the model. Technically you can give it the multiple-choice question of “which letter comes next” to make it act like a normal autoregressive LLM, but that wouldn’t really work. I started on a 4090, which was able to make decisions every 500ms, but that wasn’t really quick enough for Doom. I could probably have optimized it further but instead I just rented an H100 for ten minutes to record the demo, which got it down to a 190ms loop. I recorded the tool-calling Doom demo on the H100 too, so it’s a fair comparison. There’s an interesting research question here about how to implement choices like this, since they have to be predictable by a single token. I started with indexes but found “labels” (just picking some token to associate with the choice) performed way better on Wikiracing (though not Doom). How many choices do you have to have before labels are better than indexes? Of course, you could alter the model to directly output the choice, but I like the idea that you can do all of this in the inference code for any LLM. An every-ten-second loop that sets an overall strategic goal An every-five-second loop that sets a tactical subgoal based on (1) An every-second loop that breaks down the current tactical subgoal into specific targets A tight inner loop that runs as fast as possible (e.g. every 100ms) that controls which actual inputs are activated Technically you can give it the multiple-choice question of “which letter comes next” to make it act like a normal autoregressive LLM, but that wouldn’t really work. ↩ I started on a 4090, which was able to make decisions every 500ms, but that wasn’t really quick enough for Doom. I could probably have optimized it further but instead I just rented an H100 for ten minutes to record the demo, which got it down to a 190ms loop. I recorded the tool-calling Doom demo on the H100 too, so it’s a fair comparison. ↩ There’s an interesting research question here about how to implement choices like this, since they have to be predictable by a single token. I started with indexes but found “labels” (just picking some token to associate with the choice) performed way better on Wikiracing (though not Doom). How many choices do you have to have before labels are better than indexes? Of course, you could alter the model to directly output the choice, but I like the idea that you can do all of this in the inference code for any LLM. ↩

0 views
iDiallo Yesterday

The Front Page of the Internet Is Up for Grabs

Last week I opened Edge to test my Internet connection, because that's the only thing I use it for, and noticed something new. Right beneath the address bar was a new bar with a button offering to let the browser launch automatically whenever I start my computer. That was strange. Yes, I open the browser the moment the computer starts, but why would they want to do that for me? And just now, I opened Chrome, and saw the exact same message but for Chrome: It looks innocent enough. We all start browsing the web right away when we turn on our computers, so what's the big deal here? Well, the big deal is that everyone wants to be the front page. Defaults matter. While we all want to believe we have agency, the numbers don't lie. Most of us never change the default settings . Just last year, Google paid Apple close to $20 billion just to remain the default search engine on Apple products. While we're all distracted by the AI race, the one race that actually matters is who remains the front page of the Internet. Microsoft is trying to claim that spot by loading its browser at Windows startup. Now Google is trying to win that space back by doing the same with its own browser. On my Mac, Anthropic has been very aggressive about trying to load Claude automatically. I suspect every company will be pursuing this aggressively in the near future. The quality of your product is secondary to the exposure it gets. Not even benchmarks matter if you can't be on the user's front page. Google still holds the front page of the Internet, and that makes up for its AI model not being at the top of the benchmarks. But don't let them take this space from you. They are all fighting for your attention. I urge you to simply consider pressing the little X button they tuck into the corner. It's your computer, use it the way that suits you best.

0 views
neilzone Yesterday

Initial thoughts on the EU KIDS Act

The European Commission has proposed the EU KIDS Act . It is yet another set of Internet/web regulation proposals to examine, for jurisdictional overreach, lack of common sense in terms of material scope, and so on. It is only a legislative proposal at the moment, so it may not become law, and it may not become law in this form. Based on a quick skim, this is indeed another fine mess, full of unrealistic expectations. The proposal covers a lot of services: I have read this from the perspective of online social networking services, thinking predominantly about Mastodon and other fediverse services. At least code forges are out of scope (“open-source software-developing and-sharing platforms”). And Wikipedia seems to have its own bespoke exemption (“not-for-profit online encyclopaedias”). Small, low risk services are in scope. The covering material specifically notes: small and micro enterprises are not exempted from this Regulation, since they may equally provide harms to minors. It would undermine the objective of this proposal to exclude them from scope Wow. I wonder if the drafters will realise just how harmful this is. In terms of territorial scope, it is broader than the EU GDPR, and indeed the UK’s Online Safety Act, purporting to apply to providers of services irrespective of where they have their place of establishment where they offer those services to recipients of the service that have their place of establishment or are located in the Union I wonder if anyone working on this stopped to think about the boundaries of their laws, and whether they really think that they can impose obligations on people in other countries, merely because that person is running a service which happens to be available to people in the EU? Do I, as someone who runs my own fedi server, where people in the EU can read my toots and respond to them from their own instance, fall into scope? I do not know. Providers of online social networking services … shall not allow a natural person below the age of 15 years to create an account with that service or to access that service by means of an account, created for, or attributed to, that person, where the service poses a risk to the privacy, safety or security of a minor below that age. (Article 6(1)) The tests for “poses a risk” set an incredibly low threshold, and include: enables recipients who access the service through an account to transmit content in real-time to an indeterminate number of other recipients of the service, including through live streaming of audio-visual content enables recipients who access the service through an account to contact, communicate and otherwise interact with other recipients of the service not part of the recipient’s pre-existing connections or subscriptions So a “papers, please” web would become the norm, according to this. For example: When creating an account for a minor pursuant to paragraph 2 of this Article, the provider of online social networking services … shall take measures to establish whether the person creating the account is the holder of parental responsibility over that minor in accordance with Article 26 and verify that the recipient of the service has reached the age of 13 years in accordance with Article 28(1). (Article 6(3)) Article 26 sets out how the European Commission envisages this working, but, wow, I just don’t see it. Harking back to (what should be the exceptionalism of) broadcast regulation, there’s another banger: Providers of online social networking services… shall put in place effective measures to ensure: time-limited access for minors on their service; interruption of usage by minors on their service. Such measures shall be designed in a way that protects school time and core sleep hours of minors. (Article 9) Sorry, I have to turn off my fedi server now, because a child in a different timezone might be heading off to bed and my toots might be distracting… Some of the proposals seem to relate to core browser functionality: Providers of online social networking services … shall put in place measures to ensure that settings are set by default to a high level of privacy, security and safety of minors. To ensure compliance with this paragraph, such providers shall, by default, turn off at least the following settings: other recipients of the service shall not be able to download or take screenshots of contact, location or account information of minors or of any content uploaded or shared by minors on the service; I have no idea how the drafters of this expect the provider of a social media service available via a web browser to restrict screenshots of everything posted by a user. It is not within their gift. The only way to make this work would be either to force all access to be via an app (which would be daft), or preclude child access (which has age verification challenges). Some of the use restrictions would seem very challenging: Providers of online social networking services … shall put measures in place that ensure a high level of privacy, safety and security of minors as regards contacts between minors and other recipients of the service. Those measures shall at least ensure that: other recipients of the service are not able to initiate direct contact with the minor, if the minor has not pre-approved such contact So a 17 year old here posts something interest. No-one is able to interact with their post, unless the 17 year hold has “pre-approved” it. Oh, don’t worry, you won’t be able to see their post anyway: by default, other recipients of the service not previously accepted by the minor shall not be able to access account information of the minor or content uploaded or shared by the minor on the service online social networking services; video-sharing platform services; software application stores online games; operating systems; AI companions; general conversational chatbots. time-limited access for minors on their service; interruption of usage by minors on their service. access to microphone and camera

0 views
Unsung Yesterday

“Insert your favorite nursery rhyme.”

Make Some Noise is one of many fantastic improv comedy shows on Dropout . I noticed one of the ongoing themes is “tech gone bad,” so I compiled a short list below. I just find these really funny. Video call with shitty wifi : = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/insert-your-favorite-nursery-rhyme/yt1-play.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/insert-your-favorite-nursery-rhyme/yt1-play.1600w.avif" type="image/avif"> Logging in with 30-factor authentication : = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/insert-your-favorite-nursery-rhyme/yt2-play.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/insert-your-favorite-nursery-rhyme/yt2-play.1600w.avif" type="image/avif"> A cutscene in a video game that’s glitching out : = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/insert-your-favorite-nursery-rhyme/yt3-play.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/insert-your-favorite-nursery-rhyme/yt3-play.1600w.avif" type="image/avif"> In case you’re curious, the improv actors are: Zac Oyama, Josh Ruben, and (occasionally) Brendan Lee Mulligan. (If you’re a Dropout member, these are the original, slightly longer landscape videos inside the respective episodes: wifi , 30-factor , game .)

0 views
Martin Fowler Yesterday

I don't like LLMs

I have a lot of mixed feelings about AI and LLM technology. I’m fascinated by its effect on our profession, excited by the potential gains in productivity - and thus the products we could rapidly build. On the other hand, I’m fearful of the damage AI might cause: agent swarms taking over our virtual and physical infrastructure, designing bio weapons. But, back on my first hand, LLMs might also design miracle cures, and come up with clever ways to raise our prosperity. Fundamentally I don’t think we have a choice about riding on the AI technology train. It’s a wild ride and I just hope we’ll get through it OK. But as I mull on this more, I realize that among this mix of contrasting feelings, there is one emotion that dominates - one that comes from my direct interactions with LLMs. I don’t like them. They talk to me in this grating LLM-voice, an uncanny valley of talking to a real human. They confidently bullshit me - often giving me useful, helpful answers. But also just making stuff up with the same assurance - and with only a veneer of fake remorse when I call them out on it. That’s not enough to make me feel we should avoid them. As Jessica Kerr put it “not only are they useful, it is irresponsible not to use them…. They’re more thorough, as well as faster.” This contradictory reaction comes through in polling , where people say they find these models are useful, but also that they think they will be bad for society. Much of this may be because LLMs are young - we haven’t trained them to grow up yet. Maybe I’ll like them once they mature. (I hope we get to find out.) But I’m not encouraged when I think of the kinds of environments that cultivate them. I’m wary of the Silicon Valley brogrammer subculture, and these LLMs are their products, so naturally lean toward their world-view. When we think of AI agents, we shouldn’t anthropomorphize, treating them as conscious beings with their own will. They are (software) machines, developed by people working in corporations. While the agents’ behavior aren’t explicitly programmed, they are nurtured with the values of their creators. One of my most successful life-hacks is to avoid people I don’t like or don’t trust. I decline to interact with them socially, and make a deliberate effort to avoid working with them too, even if they are doing much that is beneficial. I feel that hanging out with pleasant, capable people, the people with integrity, has made my life a far better one. Hence my visceral dislike of interacting with an LLM that’s not just making a pretense of being human, but also posing as the kind of human I walk away from.

0 views