Note #740
Flying feels like a miracle. Thank you for using RSS. I appreciate you. Email me
Flying feels like a miracle. Thank you for using RSS. I appreciate you. Email me
Speaking of user interface guidelines, developer Matt Sephton and others compiled many of Apple’s human interface guidelines , starting from 1980, all the way to 2014, including some goodies like early drafts, NeXT, Newton, and so on. = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/i-think-theres-a-lot-of-value-in-these/1.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/i-think-theres-a-lot-of-value-in-these/1.1600w.avif" type="image/avif"> = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/i-think-theres-a-lot-of-value-in-these/2.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/i-think-theres-a-lot-of-value-in-these/2.1600w.avif" type="image/avif"> Elsewhere, designer Geof Crowl put together his own list that goes broad instead of deep, including style guides other than Apple’s, too. The list was made in 2020 and a few links are already broken, but there are some gems here like the modern Designing for Playdate , or the classic Zen of Palm from 2003. You can learn a lot just by grabbing one and scanning it. Let me add a few more I know of: = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/i-think-theres-a-lot-of-value-in-these/3.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/i-think-theres-a-lot-of-value-in-these/3.1600w.avif" type="image/avif"> = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/i-think-theres-a-lot-of-value-in-these/4.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/i-think-theres-a-lot-of-value-in-these/4.1600w.avif" type="image/avif"> #interface design #mouse #principles #style guides Amiga User Interface Style Guide (1991) Magic Cap Concepts (1995) with a strange subtitle “Everything right is wrong again” Java Look and Feel Design Guidelines (1999) Windows Interface Guidelines for Software Design (1995)
I like learning new things, and I like learning new old things. I was looking at Apple Human Interface Guidelines from 1987 and this passage caught my attention: The most common use of double-clicking is as a shortcut way to perform an action. For example, clicking twice on an icon is a faster way to open it than clicking once to select it, then choosing Open from the File menu; clicking twice on a word to select it is faster than dragging through it. I knew that double click an icon was a shortcut to the first action (typically Open), but I never really thought of double-clicking a word as a faster way to drag across to select it – even though, in hindsight, it makes perfect sense. Another vintage thing I learned of recently from a coworker is this, also covered in the 1987 HIG: If the user begins a double-click sequence, but then drags the mouse between the mouse- down and the mouse-up of the second click, the selection becomes a range of words rather than a single word. This doesn’t feel (to me) like a very pleasant gesture to perform repeatedly, but what feels nice about it is that it automatically snaps the selection to the endings of the words: Part of me would prefer this to be the default behaviour when selecting more than 3 words, or so, so you could be less precise. Anyway. The double clicking to perform default action applies to a lot of lists of things. Here are some examples from Scrivener, Word, and Lightroom – you can double click on each of these items to proceed, without having to select and click the button: = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/if-nothing-happened-you-probably-did-not-press-the-button-quickly-enough-the-second-time/2.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/if-nothing-happened-you-probably-did-not-press-the-button-quickly-enough-the-second-time/2.1600w.avif" type="image/avif"> = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/if-nothing-happened-you-probably-did-not-press-the-button-quickly-enough-the-second-time/3.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/if-nothing-happened-you-probably-did-not-press-the-button-quickly-enough-the-second-time/3.1600w.avif" type="image/avif"> = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/if-nothing-happened-you-probably-did-not-press-the-button-quickly-enough-the-second-time/4.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/if-nothing-happened-you-probably-did-not-press-the-button-quickly-enough-the-second-time/4.1600w.avif" type="image/avif"> But sometimes the creators of such dialogs forget. Here’s Screen Sharing in MacOS, and a notification in Chrome where only the slow path is available – double clicking on items doesn’t do anything: = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/if-nothing-happened-you-probably-did-not-press-the-button-quickly-enough-the-second-time/5.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/if-nothing-happened-you-probably-did-not-press-the-button-quickly-enough-the-second-time/5.1600w.avif" type="image/avif"> = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/if-nothing-happened-you-probably-did-not-press-the-button-quickly-enough-the-second-time/6.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/if-nothing-happened-you-probably-did-not-press-the-button-quickly-enough-the-second-time/6.1600w.avif" type="image/avif"> The tricky part about not being a good citizen of a shared user interface is that those omissions aren’t just local to your app – they can ruin the gesture in other places, as people’s fingers learn to distrust it in not just your app, but in general. (The quote in the title is from Apple Macintosh User’s Handbook .) #flow #mouse #text editing
Working great 👍🏻 Thanks for reading this post via RSS. RSS is ace, and so are you. ❤️ You can reply to this post by email , or leave a comment .
Hanging out at the boyhood home Thank you for using RSS. I appreciate you. Email me
Last November I replaced my oldest son's PC with an £88 iMac that was so successful that I'm now doing the same thing again for my youngest son! Thanks for reading this post via RSS. RSS is ace, and so are you. ❤️ You can reply to this post by email , or leave a comment .
Dig, if you will, a picture. It's November 1983 and throngs of people are engaged in literal fisticuffs, " dozens of people were injured, some even hospitalized ," to claim the last homely, pug-nosed Cabbage Patch Kid from the local Zayre's for Christmas shopping. In video footage, a store manager climbs a display, brandishing a baseball bat, trying to will order onto the chaos. The riots make the local, then national news, and eventually have a page on Wikipedia . It's a striking exclamation mark to punctuate 1983, with dark implications for the coming new year. Had we sunk so low? Then, at the start of December, a Christmas miracle. The hottest pop star in the world released one of the most influential music videos of all time on MTV. You know the one, where Michael turns into a werewolf, but then he's a zombie, and they do that dance, and the electronic beat "ur URRR adugga dugga dugga dugga dugga dug ur URRR adugga dugga dugga dugga dugga dug", and the Vincent Price laugh, and those crazy eyes at the end. The video for "Thriller" released simultaneous to the riots, and karmic balance was restored. Huzzah! Maybe 1984 wouldn't be so bad after all. 1984 wasn't just "not bad," it was radical. I tried to find "cutest adorable chubby-cheeked preschoolers Thriller dance" but this will do. Pop-culture nerds, in particular, would gorge themselves that year with Ghostbusters , The Terminator , Gremlins , Nightmare on Elm Street , Beverly Hills Cop , Revenge of the Nerds , Splash , Red Dawn , The NeverEnding Story , The Adventures of Buckaroo Banzai , Lynch's Dune , Karate Kid , Repo Man, This is Spinal Tap , Sixteen Candles , The Transformers toys and show, Teenage Mutant Ninja Turtles comic, Stop Making Sense concert film, Prince's Purple Rain , Do They Know It's Christmas? , Neuromancer , Elite , Karateka , King's Quest , the Soviet boycott of the Olympics , and the first MTV Video Music Awards. "Thriller" lost "Video of the Year" to "You Might Think", but did win "Best Choreography" and "Viewer's Choice." Oh, and Apple aired a commercial during Super Bowl XVIII. You know the one, with the skinheads in grey, and the oppressive fascism, and the woman in red shorts, and the sledgehammer, and that electronic siren "BEE oooh.... BEE oooh", and "1984 won't be like 1984." Apple introduced the Macintosh, a system Steve Jobs said was so easy to learn, "You can sit your grandmother in front of this computer and she’ll figure out how it works." That's an interesting way to describe a computer that internal documentation shows "we will not attempt to position the product in any way as a 'home' computer." Meanwhile, that same year, the Apple II series sold another one million units while Apple was taking its third stab at obsoleting those machines. Having shrugged off the Apple III, then the Lisa, Apple II fans would be forgiven for looking at the Macintosh's anemic launch software library (essentially just MacWrite and MacPaint ) and shrugging in agreement with Clara Peller's 1984 catch-phrase, "Where's the beef?" Tom Weishaar, writing for Open-Apple newsletter in July 1985 (a rough year for Apple), "For six years now Apple's management has been trying to build a computer better than the Apple II. In all these years, after all this money, only the Apple II has ever shown a profit. Buyers are looking for tools they can use. The Apple II has plenty of power to be useful." VisiCalc was early, clear proof that innovation isn't measured in megahertz. By 1984, developers had a lot of expertise in making the most of limited hardware. Look at the strides made in six years of software research on the Atari 2600, for example. Practice makes perfect. Rupert (later "Robert") Lissner heard the call. While the Macintosh was an interesting promise of a computing future that may or may not come to fruition, there were millions of Apple II owners with immediate needs. AppleWorks fulfilled those needs, growing a strong fanbase, spawning dedicated newsletters, lots of books, add-ons, and more. When Apple finally abandoned their 8-bit line, AppleWorks singlehandedly kept those machines productive for another decade . My only experience with AppleWorks was in the early OS X days, and that version had nothing to do with what we're looking at today. That and this only share a name; the Apple II version holds a legacy. Version 3 of AppleWorks was, according to a Compute! Magazine review, when the promise of "integration" finally paid off. Much of what AppleWorks 4 and 5 do, 3 can do with plug-ins. Version 3 is also when the program became Y2K compatible. It should be more than enough to develop a strong understanding of what made it such a favorite among the Apple II faithful. Thus far, I've only looked at one other "integrated" software package: Pipedream , on the Acorn Archimedes. It put forth the hypothesis that word processing, spreadsheets, and databases are not three separate applications, they are one and the same, chemically. While I was interested in its proposition, I was cool on its execution. It's an acquired taste. Where Pipedream blurred the line between applications, AppleWorks is very much three apps in one; a veritable turducken of productivity software. Maybe I'm a little unfair with that analogy. A turducken is a forced alliance of three fowl under, let's call it "extreme duress." AppleWorks, on the other hand, is a joyful alliance, with each individual app supporting the other two. I'd argue the word processing benefits the most, but to categorize this as merely three applications which happen to ship in one box, would miss the turkey for the duck. With 128KB of RAM and an entire suite of software loaded up, AppleWorks has 23KB remaining to do real work, as indicated by the memory counter in the bottom right of the screen. That's not a lot, but it's a bit dependent on how hardcore your work is. For example, my pre-edited text for this blog post hit 9,000 words and 43KB; the database merge (I'll talk about later) hit 24KB. Both are substantial projects, so if you're just dashing off letters, giving publishers a chance to get in the ground floor of your burgeoning career as the author of legally distinct James Bond-alike, James Epoxy, you'll be fine. Chronicling Agent Epoxy's actual adventures will likely require you to scope documents chapter-by-chapter. In the wake of the release of AppleWorks , a decent amount of third-party software was published to "add value" to this triple-threat. Using ad-hoc plug-in strategies (not originally part of the program proper), the already stuffed AppleWorks could become almost an operating system unto itself, hosting not just enhancements to the core functionality, but entirely new applications like screen savers and a MacPaint clone. Other companies took a less invasive approach to enhancing AppleWorks , offering pre-built documents ready for mail merge, recipes, phone directories, and the like. I saw a reference to a fantasy football management suite called FantasyWorks , but couldn't find any disk images for it. I suspect there's an entire world of unarchived AppleWorks enhancements sitting around on forgotten floppies. One product I did dig up, and which inspired my project for this investigation, promised to turn AppleWorks into HyperCard . When I learned of this program's existence, I had the reaction of a dog being asked, "Do you wanna go outside?" StoryWorks , written by Robert C. Moore and published by Teachers' Idea and Information Exchange, promises we can "create incredibly powerful 'hypertext' applications (stacks) (and) create 'knowledge base' stacks." That is a tall promise for an 8-bit machine, and it is one the provided example files don't quite live up to. What it can do is create "choose your own adventure" style stories. It takes a bit of boilerplate to adhere to StoryWorks's programming rules, and I will need to keep track of which "segments" (like HyperCard "cards") link to which. This sounds like perfect tasks for the database and spreadsheet to help manage. Then, I'll do some writing and formatting in the word processor, and next thing you know the potential and promise of AppleWorks's "integration" will be realized. Or so the theory goes. Navigating AppleWorks's file management can feel a little pre-Cambrian, not yet evolved into the complex, multicellular operating system you're reading this very blog within. Hardware-wise, we have the disk and the RAM. Files must be "on the desktop" to be used, which suggests some kind of "load from disk into RAM" step is necessary. That is precisely the case. First, we have to let AppleWorks know which disk is our "current disk," shown in the upper left corner of the "desktop" screen. It's akin to interfacing with a single folder in your file system, and the visual metaphor of tabbed folders is used to assist with navigating the Desktop. Once we have files on the desktop, fast switching via lets us bounce rapidly between documents making edits, while AppleWorks preserves our changes. Ready to quit? Not so fast, buddy. Sure AppleWorks appears to have preserved your work, but did you SAVE YOUR WORK ? Thus far, our changes are only committed to the Desktop and, as established when we loaded from disk into Desktop, the Desktop is not the disk. From the Desktop, we need to "Save Desktop files to disk" to really save onto disk. Earlier, I noted the "default folder" and here's where it will come into play. Saving work to disk will save everything to the default folder, regardless of its original position on disk. I wound up duplicating files into lots of fun, new locations on disk, rather than overwriting the originals, as I learned the program. But, at least my files are safe. They were saved carefully, after all. Two core tenets of AppleWorks 's approach to integration are its clipboard and keyboard shortcuts. Both work hard to bridge the natural divide between applications. The clipboard lets us move data between applications, and the keyboard shortcuts let us move muscle memory between applications. This mostly works, as when doing cut/copy/paste, text insertion/deletion, finding information, printing, and other relatively generic tools. The shared keyboard shortcuts fall apart, somewhat, in their overly ambitious attempt to enforce uniformity across all apps, even when they don't make intuitive sense. Let's look at the 'Zoom' command as a prime example. First, think about a command called 'Zoom' and envision what the result should be when invoked in a word processor context, then in a spreadsheet, and finally in a database. What did you picture in your mind's eye? What does it mean to 'Zoom'? Contrast that with to 'Find' text, which does exactly the same thing in all apps. Keyboard shortcuts aren't evenly distributed across modules, either. Printer options are available in the word processor and spreadsheet, but not the database. "Find" works across all modules, but "Replace" doesn't work in the spreadsheet. will split a window in a spreadsheet, ala VisiCalc's split window function. However, the word processor does not accept this command, though it would be a very reasonable expectation for that to work, as we saw in PaperClip on the Atari 8-bits. I certainly recognize that each program will necessarily have unique needs. However, I'm not convinced the mental fortitude required to remember how one command differs from application to application is any easier or harder than just learning a new keyboard command specific to each application. At this point, I've written thousands of words in AppleWorks and I can say confidently: the word processor is good. I have throttled the emulator to run at 1x speed, so I'm not completely head-in-the-sand about the reality of using the program on period equipment. Whatever horsepower I throw at it, it remains performant and a joy to type in. Of course, we have the usual list of modern gotchas, like the lack of international character input, and the printer-centric formatting options (which do not survive ASCII export). But the basic act of getting words on screen is smooth, along with minimal chrome which keeps us informed of the state of the writing environment. The chrome serves the document, ensuring the writer is never lost. The upper two lines show the ruler, what file we're working on, what mode we're in, and what hitting will do right now. In the above case, here in the Printer Options screen, will return me to the REVIEW/ADD/CHANGE screen. This clear explanation of what navigation buttons will do was something I appreciated about Bank Street Writer , and I'm happy to see a mature version of it in AppleWorks . The rule shows our tab stops and their respective styles, but doesn't show an indicator for the right margin. Editing tab stops is as easy as editing a line of text, thanks to fixed-width type. Just type a character for the tab styling you want at the place you want it, including centering and decimal alignment. The bottom displays a running line count and where the cursor currently sits. The bottom left has what seems like pretty useless information to me, "Type entry or use Apple commands." as a gentle reminder of how to use the program, I guess? The bottom right shows how to reach Help. That gives the UI four lines, with 20 devoted to writing. It is enough, though I do wish for a live word counter. AppleWorks 3 brings integrated spell check, with reviews of the time saying it is better than any of the third-party spell-checkers that preceded it. That's a strange note, because I find its usage convoluted. Invoked by , to "verify" the document, the "Options" it presents are obtuse and unintuitive, but all its asking is if we want to replace words one at a time or in a list, and do we want a post-verification summary of everything changed? If so, how do we want the summary presented, printed or on screen? The "Summary" is where we finally can see our word count, plus a list of all the weirdo words we used and how we corrected them individually, along with the occurrence count for each word. I'm not clear what I'm supposed to get out of knowing I used the word "turducken" 8 times. Maybe this is to help encourage the writer to mix things up? I refuse. As is typical of word processors of the day, a vast amount of the program's features are related to printer formatting and control. I do not own a printer, and "printing" to an ASCII file strips away most of the fun stuff, like bold and superscript. The spreadsheet portion is both amazing and boring. It's VisiCalc , with minor changes to meet the keyboard shortcuts (i.e. no slash menu command) and basic UI chrome of the suite. It does almost everything VisiCalc does, on a much larger worksheet and without quite as strict a memory barrier. If you have the RAM, even in this economy , you can eat VisiCalc's lunch. It also has many of the same frustrations as VisiCalc , including row-based vs. column-based calculation order, which can require a full document double-calculation to synchronize early formulas with later cell values. We're also stuck with the archaic (even by this time) "one column width for all cells," rather than per-column width adjustments. There are no graphing tools, and what options are available are so limited as to be effectively useless in a Lotus 1-2-3 world into which AppleWorks 3 was born; we'll need to rely on third party solutions to this problem. Copying complex formulas with relative cell references still requires manually selecting "relative" for every single reference in every single cell copied to a new position. Just a quick note to those cloning popular apps: you don't need to clone the terrible parts. Still, VisiCalc once turned the Apple II into a must-purchase investment and birthed an entirely new genre of business software. Now, that watershed event is collapsed into one subset of bullet points in a longer feature list on the back of the product packaging. The flat-file database module is uninspired, which is surprising to me because it is the module upon which this entire program was built. With a great word processor, and effectively a full clone of VisiCalc included, I expected similar depth from the database. Setting up fields and records, searching, sorting, and filtering, are all simple enough to accomplish. Its adherence to AppleWorks's common keyboard commands makes it pretty trivial to search for records, edit text, delete records, and print. Yet there is so much it doesn't do, I can't help but feel disappointed. To start, fields cannot be assigned value types, something even dBASE on CP/M could do years earlier. Everything is a string, without even the crudest form of data validation. AppleWorks 4 would gain the ability to at least set a field to be a number vs. text, though still no Booleans or field widths, for example. When setting up a contacts database, a common use case is to have a "notes" field to remember things like family members, follow-up discussion topics, and the like. Thanks to a limit of about 70 characters per field, no single field in AppleWorks can hold that much information. I guess the solution is to set up half a dozen "memo" fields, just in case? That kind of workaround feels quite silly to me. The remainder of the module is focused on generating "reports" and "layouts," the difference being "for printing" or "for screen." This has its place, perhaps to show a subset of data and focus on the important stuff, or to copy some data over to the word processor. Otherwise, there's just nothing to get excited about here. I can't even pretend to be excited for the purpose of writing a fun blog. Moving on! The early days of productivity took some time to figure out how best to handle object selection and manipulation. On 8-bit systems especially, modality ruled the day. Today we've settled on the paradigm, where we select an object to manipulate, then choose the manipulation. Modality-driven software was the opposite. First, choose an action, then choose the object to affect, i.e. . To move a paragraph, we first select , to indicate our intention to "move." The system switches to a text selection mode, allowing us to highlight the text we wish to move. Finally, we position the cursor at the new location for the text, hit , and the text is moved. Its backwards, relatively speaking, but easy enough to adjust to. Thanks to AppleWork s integration, we can copy/paste between any two Desktop documents. It's a little strange, because copy and paste are both considered "copy" operations, via . We copy to the clipboard from application A, then we copy from the clipboard into application B, using menu prompts along the way to specify the direction of the copy. Copying has its quirks, though. Copying between two documents of the same type behaves as expected. Across application types, copying into the word processor fares best, and most closely meets expectations. But we still encounter broken formatting, like how a line of text that spans multiple cells in the spreadsheet will break apart into discrete, tab-width sized chunks in the word processor. What we actually want is to do is "print" to the clipboard, as backward as that sounds. Depending on the module, different print options become available, to define what part of the data we want to print, and how to format it. Once done, printing to is an option. The end result is much cleaner data that is closer to WYSIWYG over the normal clipboard copy command. "Print to clipboard" is the only useful cross-application option, in my testing. As the connoisseur of modern computer software that you are, I know you know that however good something is, it could always be a little better. Most software requires us to beg and plead for the developer to make the improvements that will ease our mortal burdens. Some software embraces the notion of "extension," allowing itself to be a mere vessel for thoughts yet unthought, ideas yet unrealized. When you bought AppleWorks , it never occurred to you that you'd even want it to do outlining, until you saw ThinkTank and now you kinda wish you had something like that. AppleWorks is shockingly accommodating to our wishes. One of the core features of Lissner's engine is a crazy-efficient memory management assembly routine. This gives, what by all rights should be "not enough computer," the superpower of fast app switching. Especially with a ProDisk (hard drive) attached, AppleWorks can become essentially a graphical shell for the Apple II. A number of companies took advantage of this and created various add-on packages for AppleWorks . Beagle Bros, JEM Software, Pinpoint Publishing, and PBI Software collectively published over 100 extensions. They had a lot of competitive overlap, but provided add-ons like: This was all well and good, but there was no official plug-in architecture for AppleWorks ; it was every developer for themselves. Installation routines were effectively "patches" to the base application, and there was no guarantee that patch A wouldn't step on patch B's toes during the installation process. It put some burden on the end-user to detangle things like installation order, or even compatibility with other add-ons of interest. Beagle Bros wanted to solve that problem once and for all. TimeOut was their solution to a simple, universal plug-in architecture for AppleWorks . Developers targeting TimeOut could rely on it to do the down-and-dirty interfacing with AppleWorks's memory management routines, without having to worry about patching or memory collisions. For the end-user, adding new TimeOut modules to their system was as simple as copying a single file to the AppleWorks directory. This worked like a champ, and eventually led to Beagle Bros being the contractor for AppleWorks 3 . As the main app developer, they could steer its growth to align with their own vision, and so TimeOut became part of the foundation of AppleWorks proper. This then made it super simple to produce updates of significant value to the end-user, because those upgrades already existed in the form of TimeOut extensions. And so, AppleWorks 3 received its first built-in spell checker as a result. As time went on, the outliner, expanded Desktop, better clipboards, and the like all became built-in features to AppleWorks . The big one, from my perspective, is TimeOut UltraMacros. I have it installed as an extension here in AppleWorks 3 , and in AppleWorks 5 it became a built-in feature. With it, we get a kind of mini-programming language, which includes if/then/else statements, variables, loops, value comparisons, and more than enough to fill out a 110-page manual. Macros work across all AppleWorks apps, and since the language is identical across apps, if you know how to script one, you know how to script all of them. Still a few more years until the 6502 reaches its maximum potential. With the powerful combination of AppleWorks + UltraMacros + StoryWorks , we have a veritable turducken of creative utility. In later writeups espousing the benefits of the AppleWorks plug-in system, it was suggested that future application authors would be writing exclusively with AppleWorks as the hosting application in mind. I can see the appeal. Unfortunately, StoryWorks didn't get that memo, so the workflow between AppleWorks and StoryWorks isn't the seamless experience I would have liked. They are two completely separate programs; quit one to enter the other for each phase of the writing/debugging development cycle. Let's all take a moment to re-appreciate our multitasking operating systems. Now, I've never written a Choose Your Own Adventure (CYOA) before. In browsing the books of my youth, I can see how AppleWorks's integration could be useful in designing such a thing. I can keep multiple documents open simultaneously, and jump quickly between them with . Here's my plan. With the spreadsheet, I'll keep a list of page branches. This will sketch out the core content on each page, with page numbers showing where each choice leads. In the database, I'll set up a template for the "programming" information for each page, based on the spreadsheet sketch. The database's 70-character limit on fields means I can't type the entire story into the database, but I can rough out a plot point or two. Then, in the word processor, I'll merge the database into a template to save me from tedious, manual boilerplate formatting. Once merged, if I've done my job right, that should build a skeletal game I can dry run through StoryWorks . "It's so simple, I can't believe people struggle to write games," the author scoffed. OK, look, I was out of pocket and I take it all back. In the previous section I was a naive child, unaware of his own ignorance. That was "Then Me" and he was a fool. "Now Me" is wiser, less prone to shooting his mouth off about things he doesn't understand. Making games is hard, I can admit that now. My biggest stumbling block is conceptually very simple: I don't know how to write a CYOA. What seemed like a clear plan in my mind turns out to not be that in practice. I thought, "Start with a rough sketch, start filling in details, repeat until done." would work. So, fine, all this really means is that I'm not a game designer, a fact I actually already knew. Looking through websites about how to create such games, everyone seems to have a different take. Spreadsheets as an organizational tool comes up a lot, as does sketching out decision trees, for which the "Paint" add-on might work, but I have a better plan. For my first attempt, I simply don't have the wherewithal to create something this intricate. My goal is to push the software, not my brain, so I'm going to be a thief and steal someone else's idea. Or rather, I'm going to steal someone else's structure. The original Choose Your Own Adventure series, the "fourth best selling children's series of all time," was established by Edward Packard in 1976. Relaunched by Chooseco in 2010 , the classics have been reprinted in handy box sets, and even new adventures have been published. If you've ever wished for a Cthulhu-themed CYOA, aimed at readers 9 - 12, your very specific prayer has been answered. The original CYOA books are also available on the Internet Archive's "Open Library" program. My intention was to get "inspired by" those, but instead I'm going to swipe the basic decision tree from one, and write my own story. There is a spy adventure called "The Deadly Shadow," by Richard Brightfield, that I will lean on for the heavy lifting of providing the structure for my story about low-rent super agent, James Epoxy. Ah, let's just go ahead and call it a parody at this point. There's no point in kidding myself that I'm about to do anything original. In working through my failed attempts, I encountered some frustrating limitations of AppleWorks integration. In AppleWorks , we have an option to set numbered "markers" throughout a word processing document. These are invisible tags that can be jumped to, by number, for quick navigation to known locations. StoryWorks relies on these to delineate story "segments." Just as AppleWorks can jump to markers, so too does StoryWorks use the same principle to jump to segments. The difference is that StoryWorks will only display text up to the next segment marker, effectively splitting the text into something akin to HyperCard "cards." And so, like HyperCard , StoryWorks also calls a collection of cards a "stack." In setting up my word processing template, I added segment markers where appropriate. Merging in the test data went smoothly, the result of which was a new document written to disk as ASCII. Maybe you already smell the trouble? ASCII conversion stripped out all non-printable tags and markers, rendering it StoryWorks- incompatible. That means adding StoryWorks markers must be done post data merge. Luckily, I have UltraMacros installed, so complex, repetitive, menu-driven marker creation is a simple self-defined keystroke away. In fact, because macros are just "a list of keyboard commands performed in order" I should be able to semi-automate find-and-replace placeholder text of my own design with a real StoryWorks marker. If I run that a few times I can prep my document for StoryWorks ingestion lickety-split. Does anyone still use that term? Where did that phrase come from, anyway? One thing I've quickly learned is that debugging is tricky. StoryWorks will parse the AppleWorks document and list its errors, but in practice I found that an early error can cascade, generating pages of errors. All we can really do is fix the first one and test again, hoping the later ones will disappear. So the proper strategy is to test early, test often. Before there is anything resembling a cohesive narrative, we need to make sure the skeleton of the project, as merged from the database, compiles and works. For this purpose, it actually does help to have even truncated page descriptions populated from the database, just to get a yes/no understanding if navigation is working. First, here's the template I've come up with. and placeholders in the template align with field names in the database. During a merge, possible field names are presented in a list for insertion, making it impossible to accidentally mistype one. means "don't insert a blank line if this database entry is blank." will always insert something , even if there is no entry in a record. Here's what a representative record looks like. You can see the field names are referenced in the template, above. As I mentioned earlier, the merge gives us an ASCII text file. That file, before adding true markers via macros, looks like this; note the text I use as a placeholder for where the macro should place a real page marker. I see that the mail merge did not follow through on the promise of skipping blank entries. I also see that, despite showing me centered text on screen in my template, that was stripped from the document as well. I have quite a bit of cleanup to do to tighten up these layouts. This is becoming work. Next, I'll run the macro I created to automate marker insertion. Macros are action-for-action transcripts of the keyboard actions required to accomplish a task, including cursor repositioning or text cleanup to prepare for the next step. Simple macros can be built simply, because macros are defined in the language of the AppleWorks user . I cannot praise this approach to macro scripting enough, and I will bring it up every opportunity I can. The more I encounter it, the more I honestly feel something fundamental has been taken from users over time. It's one thing to allow someone to record their actions blindly, with a slightly patronizing "There, there, don't worry your sweet little head about how this all works." It's quite another thing to tell a user, "Hey, you know all those tools you've been learning? You can use those exact same skills in powerful new ways." It would be relatively simple then to teach someone how to wrap an existing macro in a loop or decision construct ( UltraMacros can do these things and more), building up a foundation of self-confidence. It must surely be far more accessible than whatever Google is proposing. I mean, speaking of Cthulhu! After running my macro to the end of the document (holding down the macro shortcut auto-repeated until finished), here's the final result. I have to to "zoom" into the document and reveal the hidden formatting codes, but so far so good. Everything appears to be ready for StoryWorks . Fingers crossed. Hot dang, got it on the first try! (as far as you know) Now that the basic structure seems to be in working order, the last thing to do is write an entire novel. *dry cough * What a delightful piece of software; a delicious turducken! It will make excellent sandwiches for tomorrow's lunch. Easy to learn, easy to use, and even relatively complex cross-application actions are achievable with a gentle learning curve. I found I only needed manuals and books for very rare, "Can AppleWorks even do this?" questions, and very rarely for, "I'm stumped." Of the word processors I've covered to date it's easily my favorite. Typing is responsive and editing is intuitive. I like the advanced tools, like markers, though I think other tools could be made easier to use or expanded. Keyboard shortcuts became second-nature very rapidly. Tasks I usually dread, like mail merge, are trivially accomplished. The database " is ." It exists. It's fine. Perhaps its simplicity is a virtue, to some degree? I'm not personally smitten, but I'm glad to have it, and it did solve a real problem for me with the CYOA construction. I enjoyed the spreadsheet as much as I did VisiCalc , which is to say it's very good but Lotus 1-2-3 still wins the day. Still, it's a lot of bang for your buck, considering it's only 1/3 of this package. Being able to easily share data with the other modules elevates its usefulness over VisiCalc . Its convenience gives it a clear win. Plus, it didn't include "Copilot integration" long before Microsoft removed it from Excel . Perhaps the biggest takeaway for me is the stark reminder of how much can be achieved with so little: so little CPU, so little RAM, so little hard disk space. One can't help but ponder, "If this is what 128KB can do, imagine applying the same discipline toward 1MB of RAM." Writing for inCider , Oct. 1989, Senior Editor Paul Statt compared the development cycle of Lotus 1-2-3 Release 3 (a rather notorious, oft-delayed "update") with AppleWorks 3 . Both were initially created by lone developers, Jonathan Sachs and Rupert Lissner. Release 3 of Lotus took a team of 40 developers years to write and shipped on 14 floppy discs. AppleWorks 3 was written by three Beagle Bros developers and shipped on two double-sided discs. An update to Lotus likely meant also updating one's computer to a 286 with 1MB of RAM and a hard drive. AppleWorks 3 ran on the exact same 128KB machine the version 1 ran on. Almost 40 years later, his complaint feels uncomfortably modern, don't you think? In recent news we read of ways to whittle Microsoft's 1GB weather app "down to" 130MB RAM consumption. While on the other hand, we have something like Weatherbot , occupying 6MB RAM on Windows 11 and about 2MB on System 7.5.5 (yes, it runs natively on both and more!) The call for an efficient use of system resources is clearly still championed by some, but I don't see much of a call for brutal efficiency. Using 1/100 the RAM of a Microsoft-made app is fantastic, don't get me wrong, but let's get far more ambitious. No more thinking in megabytes, think in kilobytes ; express ambition by orders of magnitude . AppleWorks singlehandedly kept Apple IIs in production use for decades , well into the GUI era, well past their "prime." We deserve such longevity again. We need it. There is a financial cudgel of planned obsolescence that beats us down for lunch money on a regular basis. We're told it's for progress, but that proposition holds no tether to reality when we can see and touch a 128KB rebuttal that proves the bully a liar. Don't worry, I wouldn't leave you hanging. The beginning of my James Epoxy CYOA; pause to read. The description of Dimitrius is straight from the source book; I did not set that up for my running gag! Ways to improve the experience, notable deficiencies, workarounds, and notes about incorporating the software into modern workflows (if possible). As with many tools of this type and era, the program's emphasis on "printing" limits our formatting options pretty drastically. We can only embed non-printing printer codes into a document, and that requires setting toggles for (say) bold to start and stop at specific points in the page. It's anachronistic in all of the non-nostalgic, annoying ways. Late in my review cycle I came across a project keeping ProDOS alive on real Apple II hardware. The last official version of Apple ProDOS was 2.0.3 in 1993, but this project is at 2.4.3 with 2.5 on the way. John Brooks has been maintaining this for years now. It includes a kind of app fast launcher called Bitsy Bye which might smooth the process of switching between apps (like between AppleWorks and StoryWorks , in my case), if for no other reason than it appears to eliminate keystrokes and simplify file navigation. 1984 rocked. AppleWin x64 1.32.00 on Windows 11 Emulating a prohibitively expensive Enhanced Apple IIe. All cards and hard drive included, maybe $8,000? ($22K in 2026) 3x machine speed and enhanced disk access Super Serial Card Mockingboard C Disk II w/two floppy drives Hard Disk Controller with 5MB ProDrive Z80 SoftCard RamWorks III (3MB) AppleWorks 3 TimeOut UltraMacros In the database, will "zoom in" and "zoom out" between the record list and an individual record. In the spreadsheet, it will toggle display of raw values or formulas embedded in the cells. In the word processor, it will reveal hidden formatting, like for printing, navigation tags, bold/italic format codes, carriage returns, and so on. Graphing and plotting spreadsheet data, for the Lotus -envious Expanding the limits of open documents, clipboard, database size A full clone of MacPaint Appointment calendar High-resolution font support The new version of AppleWin x64 makes it super simple to install a whole host of fun circuit boards into various slots. You can build the pimped out Apple IIe of your dreams effortlessly. I mostly ran at 300% CPU speed and experienced no quirks, repeating keys, crashes, or anything else unbecoming of a well-behaved computer. AppleWorks never crashed; the recent update to AppleWin x64 worked perfectly. I did have AppleWorks fail to bring in some fields when I copied from the database into the word processor. CiderPress2 works perfectly for opening Apple II disk images. Hard drive images are of file extension, and CiderPress2 can manage these, including copying file structures between disk images to cobble together your own custom disk from others. CiderPress2 can natively open AppleWorks documents to show formatted content. Personally, I found it better to print from AppleWorks to an ASCII file, then copy out the text from that via CiderPress2 . Doing so will ensure no hidden AppleWorks formatting codes are copied over. Modality The modal nature of the editing tools limits the power of macros. If we could select some text, then apply a macro to that selection, that would be fantastic. Markdown keyboard shortcuts would be a breeze! StoryWorks The biggest issue I have is how there is no option for exporting standalone stories. Being able to build a bootable Apple II floppy would turn it into a great, simple game maker; translations of existing CYOA adventures would almost be self-coding. I would also like for it to respect formatting codes, like centering. Spreadsheet Though opening DIF files is an option, I had zero luck getting it to open my VisiCalc DIF exports. It doesn't match the integration goals of the program, but I did miss the slash menu; it could have been a nice alternate UI for those transitioning to AppleWorks . Editing cells, such as to change a cell from a label to a value, always tripped me up. Considering this is version 3, well after Lotus 1-2-3 hit the scene, I do wish for better control over column widths and some concession to graph making. 3D spreadsheets, linking a cell in one sheet to an entirely different file, would be useful; this feature debuted in AppleWorks 5 . Database Basic concerns about being unable to apply value types to fields, or any kind of data input validation, are addressed in AppleWorks 4 and 5. Form layout formatting tools are overly simplistic. Being able to merge into the word processor directly from disk, rather than from the clipboard, would be helpful. Word Processor I don't have much to complain about. It is more robust than it first appears, with a gentle learning curve. I think word count should be surfaced to the main UI, and the on-screen ruler could be more informative.
This is an external post of mine. Click here if you are not redirected.
People are pretty unhappy about Anthropic’s recent announcement that they’re planning to include a hidden watermark in Claude model outputs. Will this lead to a mass exodus from Anthropic models? Will the introduction of watermarking be a meaningful change for users? No. AI text watermarking is not a big deal. It doesn’t make the text worse, it doesn’t make AI outputs more detectable in practice, it doesn’t violate user privacy, and everyone’s going to be doing it by 2027 regardless. There is no meaningful difference in quality between watermarked and unwatermarked text. I wrote about this more here , but the two popular ways to do it — Google’s SynthID-Text and Meta’s TextSeal — are completely transparent to the user. They work by replacing the pseudo-random logit sampler with a different pseudo-random logit sampler. Suppose you were gambling on coin flips with your friends, and instead of flipping a coin you decided to do this: That would still be random enough to gamble with, right? But, like a watermark, you could theoretically go back and identify that that method was used, so long as you recorded the exact time of each “coin flip”. Text watermarking works the same way: it chooses a method of “randomness” that can be detected after-the-fact. Watermarked models will not be any less capable than unwatermarked models. What about cases where the model is quoting something, or giving you the answer to a mathematical problem, or doing something else where the output is largely pre-determined? Wouldn’t enforcing a watermark there make the output worse? It would, which is why none of the AI labs are going to do that. Text watermarking approaches only replace the existing randomness in the logit sampler: in any case where the model is always going to pick the same tokens, there’s basically no randomness to play with, so there won’t be a detectable watermark in those tokens. I think all this comes from a worry that you were previously getting the best token, but now you’re getting a lower-quality token that satisfies the watermark. For instance, Anthropic’s announcement suggested that the watermarking is visible in choices like the decision between “overcast” and “grey”. Many people have predictably come out to say that decisions like these are really important to good writing, and that only an illiterate tech bro could think these words are identical. This is a misunderstanding of Anthropic’s position and of how watermarking works. Specifically, it’s a misunderstanding because it suggests that the unwatermarked model would choose “overcast” while the watermarked one would choose “grey”. This is not how it works! If Claude Fable prefers “overcast” to “grey” in a particular context (say, 80% to 20%), you’ll get “grey” 20% of the time from both the watermarked and unwatermarked model. Models already include a healthy amount of randomness in order to promote creativity. Text watermarking just introduces a way to make those random choices that’s detectable after the fact. The other big reason to not worry about AI watermarking is that AI text content has always effectively been watermarked . Most careful readers can tell when they’re reading AI outputs , because language models tend to gravitate towards certain habits of language : em-dashes, rhetorical opposition, punchy one-liners, “claudese”, and so on. In fact, it’s possible to train classifier models that reliably distinguish AI from human writing. From what I can tell, some of the backlash to watermarking comes from people who buy AI inference in order to pass it off as their own work, and who worry that watermarking will make it harder for them to do that. For these people, the watermarking announcement is akin to Anthropic saying “hey, instead of making you seem smart, we’re going to publicly brand you as AI users and make you seem dumb”. But of course this has always been the case! Nobody who is currently getting away with passing off AI outputs as their own will be caught by watermarking. For the majority of cases, it’s already painfully clear what’s happening for anyone who reads the slop . For sophisticated AI users who are avoiding the “house style”, any suspicious readers who would paste their stuff into Anthropic’s watermark detector could already have been pasting it into Pangram . Tools like Pangram 2 only give you an estimate of the chance that output is AI-generated. Wouldn’t a watermark be a more solid confirmation? Not really. Text watermarks are probabilistic too, because any token chosen by SynthID could theoretically have been chosen by a human. I suppose the Anthropic watermark page could be considered more trustworthy than Pangram, because it comes right from the source, but it’s not impossible that in some cases Pangram might actually be better at identifying AI-generated text than the watermarking too. I’ve also seen theories floating around that watermarking encodes secret content into your outputs, or somehow tags outputs with your personal information. I don’t think AI labs are using watermarks to encode data into your outputs. Text watermarking is hard : like I just said, you can’t do it when the model can only respond with the same words, it doesn’t work for very short responses, and even on long responses it can only provide a probabilistic fingerprint. And that’s encoding one single bit 3 of information! I’m not saying that encoding longer messages into a watermark is technically impossible — there are papers describing ways it might work — but there’s no way any of the labs are doing it 4 . If they wanted to associate you with your responses that badly, they’d just secretly store every model response they generated. Another reason to not get too angry at any individual AI lab for watermarking is that every single AI lab is going to do text watermarking this year . It won’t just be Anthropic. The alternative is to completely stop doing business in the EU, because of the EU AI Act . That’s currently a sixty-billion-dollar market. I am not a lawyer, but to me it seems genuinely unclear whether an AI lab could even legally do something like only watermarking EU responses: short of having an entirely different service, the plain text of the Act seems like it applies to any service offered in the EU , not just the content that service outputs to EU citizens specifically. If people really hate watermarking enough, some labs might stand up a completely separate EU service, or make an aggressive interpretation of the EU AI Act and see how the legal battle goes. When I try to be maximally charitable to anti-watermarking histrionics, I adopt an interpretation like this: people are saying that watermarking is an invasion of privacy and makes outputs worse and so on not because they believe it, but because they’re trying to pressure AI labs to firewall EU AI regulations behind a completely separate interface. In this case, it probably doesn’t matter — text watermarking is not a big deal — but I can see an American consumer being worried about more aggressive future regulation, and wanting to draw a firm line in the sand as early as possible. Interestingly, this might be very slightly even-favored. I am not sponsored by Pangram. As I understand it, Pangram is by far the best AI-detection tool right now (in part because many of its competitors are shady and exist to promote paid “AI-detection-evasion” services). Technically, this is called a “zero-bit watermark”, because you can’t recover a yes-or-no value from the watermark itself (merely from the presence of a watermark). I saw someone suggesting that an AI lab could use per-user secret keys to watermark text, and then simply iterate over the keys to figure out who generated what. I just don’t see how you could do this at any scale: watermark detection is cheaper than model inference, but it’s still (a) computationally intensive enough to be implausible, and (b) probably has a high enough false-positive rate that any run against hundreds of millions of users would match multiple people. Check the current time since midnight in seconds Count that many words forward in the Encyclopaedia Britannica Count whether the word you land on has an even or odd number of letters 1 Interestingly, this might be very slightly even-favored. ↩ I am not sponsored by Pangram. As I understand it, Pangram is by far the best AI-detection tool right now (in part because many of its competitors are shady and exist to promote paid “AI-detection-evasion” services). ↩ Technically, this is called a “zero-bit watermark”, because you can’t recover a yes-or-no value from the watermark itself (merely from the presence of a watermark). ↩ I saw someone suggesting that an AI lab could use per-user secret keys to watermark text, and then simply iterate over the keys to figure out who generated what. I just don’t see how you could do this at any scale: watermark detection is cheaper than model inference, but it’s still (a) computationally intensive enough to be implausible, and (b) probably has a high enough false-positive rate that any run against hundreds of millions of users would match multiple people. ↩
Historically, I've only ever known two great ways of focusing Emacs windows: the built-in command, which cycles through all the available windows, and the ace-window package, offering random window access. It's not that there aren't more or better ways, I just didn't look any further as these were enough for my needs. While I really wanted to make my default choice, for whatever reason, it never stuck. with a custom binding always felt like the smoother fit for my limited needs. You see, I hardly ever have more than two visible windows, so whenever kicked into action (for 3 or more windows), it often took me by surprise. The one area didn't fit my needs revolved around focus requiring visual feedback, but I eventually solved that with winpulse (a little package I wrote). As you can see, I'm a simple man using few Emacs windows, and when it comes to frames, I almost never use more than one. That is until somewhat recently, when I built ytr , a tiny YouTube radio player that sits in the corner of your frame. While it all feels fairly integrated into your frame, under the hood, renders in a separate frame. This broke my trusty flow. I couldn't just focus my radio window using my well-internalized binding. Turns out, actually caters for focusing windows across frames, but only when invoked programmatically. Sure, I can wrap it with my own custom command, but Emacs already had me covered. I found the built-in command. All I had to do was bind it to and Bob's your uncle . I can now switch between my current window and my YouTube radio, with my dear binding. Balance restored.
To get promoted as a software engineer in Big Tech, you generally need to show “ownership” of a product or area. This means being responsible for it and deciding what should happen over time. I often hear more junior engineers ask: “How can I demonstrate ownership when my manager hasn’t given me anything to own?” This way of thinking is a trap: an easy one to fall into, but a trap nonetheless. The problem is that they are treating responsibility as a complete, self-contained opportunity that someone must give them first. In my experience, they’ve got it the wrong way round: responsibility is demonstrated before it is formally given. Put yourself in the shoes of your technical lead (TL): by “giving” you responsibility, they are implicitly accepting the consequences of your decisions. They need to trust you not only to get work done but to exercise judgement without constant supervision. The most important question they ask themselves is: “Does this person deal with problems the same way I would, or better ?” If the answer is yes, they can trust you with more responsibility. When I started working on Perfetto , I built its trace-analysis tooling mostly under my TL’s direction. As I began making implementation decisions myself, he would often spot major issues I had missed. For example, I built separate classes for three or four data structures I thought were unrelated. My TL insisted they were all tables, even though I couldn’t see how to unify them. Today, Perfetto’s trace processor has more than 100 tables, all declaratively defined and generated from that common abstraction. I always tried to understand his thought process: how had he even thought of something I had missed? What approach had he used to get there? Over time, I adopted his approaches myself, using them to pressure-test API changes or performance improvements before proposing them. More of my ideas started receiving a simple “yes, go ahead.” By then, I was setting the agenda myself: talking to users, understanding their problems, translating those into code changes and deciding what to improve next. Sometimes the dynamic even reversed: he would suggest something, and I would explain why it would not work. Eventually, my TL started calling me “the owner of the trace-analysis tools” instead of “the engineer who works on them.” To be clear, this does not mean grabbing projects or stepping on other people to demonstrate ownership. If you exercise sound judgement within your existing scope, a reasonable TL should gradually trust you with more responsibility. Of course, this does not work with a bad manager or in a bad environment, where earning responsibility may be more about politics than judgement. Formal responsibility is a trailing indicator, not a leading one. You’ll find that people begin trusting and deferring to your judgement before the responsibility becomes explicit. Responsibility is taken in small increments before it is given in full.
This is part 7 in a series of posts on writing concurrent network servers. In this part, we discuss how the challenges described in earlier parts are tackled in the Rust programming language. All posts in the series: Several years have passed since the previous parts were published. I've recently went over them to make sure the information presented is still relevant and all the code samples build and run using modern toolchains. I strongly recommend reviewing the previous parts before reading this one. This post assumes a basic familiarity with the Rust programming language. It will only explain Rust constructs when we encounter code that wouldn't appear in an introductory book or tutorial. The first few parts in the series focused on a socket server that implements a simple state machine protocol. See part 1 for a complete description of the protocol. Let's start by showing how this protocol is implemented in a basic sequential Rust server: With the function serve_connection defined as: As a reminder, this server version is sequential because it accepts clients one by one; the main loop blocks on serve_connection until it's done (the client closes the connection), and only then goes back to accept the next client. Clearly, handling clients one by one won't do. In part 2 , we've discussed approaches that use OS threads to handle clients concurrently. Let's start with the unbounded one-thread-per-client solution in Rust: The spawn method returns a Result<JoinHandle<T>> ; on success, we allow the handle to be dropped at the end of the loop iteration. In Rust, this detaches the thread; we don't actually wait for it to complete. This is reasonable for our code sample, because the loop is infinite ; it never terminates anyway. The potential for runaway threads is just one of the issues with the unbounded threads approach discussed in part 2. The solution is to use a fixed thread pool. Before diving into the code, a quick note on the design: the thread pool is a fixed set of threads that await "jobs" and handle them to completion. In our case a "job" is serve_connection for a specific client. There are many ways to implement a thread pool; for our use case, I went with a set of threads that all get a shared channel to which the main thread sends jobs. A worker thread picks up the next job from the channel, serves it to completion, and goes back to waiting for the next job. Here's how this looks in code: What is Receiver ? It's a type from the crossbeam_channel crate: Rust's builtin channels in std are mpsc - multi producer, single consumer, but what we need for our job queue is a channel that supports multiple consumers (the worker threads). While std does have mpmc , this is an experimental API only available in nightly versions at the time of writing. Therefore, I've opted to include the crossbeam_channel crate that provides well-tested mpmc channels for this sample [1] . And here's the main function: Note that our job channel is bounded - it has a fixed size. This helps naturally implement a backpressure mechanism - if too many clients connect, the following clients will have to wait - the main loop blocks on tx.send and won't accept additional clients on the socket until jobs are cleared from the channel. In parts 4, 5 and 6 of the series we've discussed event-driven , or asynchronous servers. Let's see how it's done in Rust. Specifically, part 6 presented a gradation from callbacks to promises to async/await mechanisms; Rust supports all of these and - as you'd expect - modern code is usually written with async/await while hiding all the details of promises (called futures in Rust) underneath. Without further ado, here's our simple state machine protocol in asynchronous Rust: Rust takes an interesting approach to async programming: it supports some of its fundamental building blocks (like futures and the async and await keywords) in the core language, but leaves the actual async engine implementation (the thing that implements the event loop) to external crates. By far the most popular crate for async programming in Rust is is tokio , so that's what we're using here. After reading the JS code in part 6, the Rust snippet above should appear fairly familiar, except perhaps the explicit tokio task "spawn". Instead of enqueuing a callback on the connection returned by listener.accept , the code spawns a tokio task, which can be seen as a green thread , and hence uses similar terminology [2] . These tasks must not issue blocking calls; therefore, they are supposed to use tokio's I/O utilities instead of the usual, blocking std utilities. In fact, we have to implement an async version of serve_connection to make this work: Note how similar this code is to serve_connection from earlier; the only real differences are the await calls on socket reads and writes [3] , and the types involved. For example, instead of a std::net::TcpStream used in the synchronous samples, here we're using tokio::net::TcpStream . Tokio has an underlying dependency called mio to handle non-blocking APIs for all kinds of I/O. It wraps OS-specific event loops like epoll to do so efficiently. While most of the series has been using a simple state machine server as the driving example, part 6 switched focus to a server for primality testing which simulates long compute tasks. Let's see how this is done in Rust with tokio: This code is very similar to the previous snippet conceptually; isprime is: Note that this sample demonstrates a job that can block (simulated with a sleep in this case). This can be problematic in an async context, as the tokio documentation explains . One potential solution would be to dispatch a blocking task to a separate thread pool and use tokio channels to communicate with it; this is similar to the approach we've taken in the thread pool sample above. Part 6 also included a version of this server that caches data on a local Redis instance; the goal was to demonstrate the complexity of event-driven code when additional layers of callbacks are added and how async/await can help mitigate that. Here's our Rust version of this server, using the redis crate (that has a tokio component enabled explicitly to support async calls): In conclusion, while Rust provides excellent support for async programming, it doesn't solve its inherent issues like function colors and the need for careful separation between blocking and non-blocking tasks. These issues are typically surmountable with some extra care, and async programming with Tokio in Rust is very popular due to its performance benefits. All the code for this post is available on GitHub . Part 1 - Introduction Part 2 - Threads Part 3 - Event-driven Part 4 - libuv Part 5 - Redis case study Part 6 - Callbacks, Promises and async/await Part 7 - Rust (this part) Because of the function color problem , the redis crate has a connection constructor specifically for async: get_multiplexed_async_connection . Here we have an example of shared state between tokio tasks - the Redis connection. Note that we don't require any particular synchronization because MultiplexedConnection is Clone ; cloning it to different tasks is safe - and in fact that's what we do for each new task. There's no magic here; if you look inside MultiplexedConnection , you'll see that it already has all the synchronization mechanisms implemented internally, as needed. Due to the magic of async/await, the code in serve_client is nice and linear. We simply await on the Redis call, and once it's back we continue with the rest of the handler. Since we're using an async Redis connection, in case waiting is required, control will be ceded to some other task that's not currently blocked on I/O.
Substack recently launched its AI detector feature in the UI, which is super interesting. Separately, lots of people asked me about interesting local do-it-yourself LLM projects as demos to show what small language models (SLMs) are capable of. Putting one and one together, I thought it would be interesting to show how an AI detector can be implemented. I will also use it as a verifier to train a small language model to produce text that avoids detection. This is a small educational project for studying the limitations of AI detectors and exploring a verifier-based LLM application beyond regular reasoning models trained on math and code. Figure 1: Substack now features a built-in AI detector. So, as mentioned above, the intended goal of this tutorial is to explain how AI detectors work by building (a simple) one. In practice, such a detector can be used to filter out spammy content, but also to potentially improve your personal writing without turning it into AI-generated text. For example, if you wrote a lengthy article and want to improve spelling and grammar, it is tempting (and actually useful) to use a grammar checker to polish it and improve readability. There are different services for that, including general-purpose LLMs like ChatGPT. However, this also runs the risk that these tools turn your writing, even though it’s still your own writing, into something that is then overpolished and now sounds like AI and gets flagged as spammy content. For example, with an AI checker, one could say, “Fix my grammar while ensuring that my text still scores 0% AI-generated.” Anyway, while we are building a fully functional checker here, the goal is to explain 1) how AI checkers (can) work and 2) use this as a case study for a more general topic on how to build a scorer or verifier that can be used with LLMs. Disclaimer: AI checkers are essentially a cat-and-mouse game. AI checkers may learn to detect a certain pattern that is indicative of AI-generated content. Then, the next LLM may incidentally or deliberately not exhibit that pattern and avoid detection. The AI checker then has to be updated to detect said LLM, and so forth. Plus, it’s also likely to encounter false positives (human written text flagged as AI-generated), but more on that later. There are several goals of this project. The overarching goal is, of course, to illustrate how AI detectors work and show an applied end-to-end LLM project including evaluation, training, and local deployment for real-world use. The outcome of this is an AI-detector API that can be used by humans and agents, and a user-friendly UI. Figure 2: Preview of the local browser interface developed later in this project. It returns a whole-text AI score and can also highlight the scores for individual text chunks. Here, we are going to develop a method similar to Pangram models, which, as far as I know, are behind Substack AI detection feature. I wrote a short article about AI-text detection a while back in 2023: What Are the Different Approaches for Detecting Content Generated by LLMs Such As ChatGPT? And How Do They Work and Differ? In essence, there are different ways to detect AI-written text, from supervised classifiers and perturbation-based probability tests to perplexity measures and watermarking. In this tutorial, we will build a model that returns a 0-100 score. It’s essentially a classifier with an estimated probability score. The probability score will denote how likely a text is AI-generated according to the classifier. (Or, to be precise the score is the classifier’s estimated probability for the AI-generated class based on its training distribution. However, we shouldn’t interpreted it as a general probability that the text was written by AI.) For this, we are going to fine-tune a DistilBERT classifier (similar to what I described in one of my early Substack articles, Finetuning Large Language Models ), but more details on that later when we get to that stage.
I’ve been playing Spelling Bee for a few weeks now, where you guess words from a selection of seven letters, one of which must be used in every word. The letters are arranged in a hexagon with that primary letter in the middle. If you want to learn more about how to play it, follow the directions here : For information on how to play, select the More in the top right corner of the game and then select How to Play . It has surprised me. I have learned that “oleo” is a word and that you can tell how hard the puzzle will be by looking at the rankings. But what I really wanted to talk about is the lessons I’ve learned from playing and how they can apply to your career. Switch Things Up There is a button that can flip around the arrangement of the letters around the hexagon. The primary letter remains in the center, but all the others are moved around randomly. When I first saw that, I thought it was silly. But I found it to be super useful over time. When I’m looking for words, I get stuck. Pressing the rotate button reveals new words because it changes my perspective. Just moving the letters around makes new words visible to me. In real life, when you feel stuck, you can sometimes make progress by flipping things around. I’ve done this in my career. I quit my first head of engineering job where I was trusted by the CEO, had loads of autonomy and worked on software that helped people find homes they loved. After, I enjoyed the freedom of lucrative contracting and learned a ton about other tech stacks. I took a quarter life sabbatical when I was around 25. After months of leisure, I learned that, yes, I really did like reading about parsing XML. I left the food-focused startup I’d co-founded (an area I’d dreamed of working in) which had just raised a seed round. Departing led me to an entirely different area of the software industry that I love and never would have explored at the startup. There are of course other ways to gain new perspective than quitting jobs: changing routines, asking friends for advice, or even taking a new route to work. Before you give up on something, flip the script and see what the results are. Just Try Something Sometimes you just have to try things. I will often put in words that I know are not valid, just mashing different combinations of vowels and consonants, like “manu”. This is not an English word, nor will it ever be. Spelling Bee also disallows proper nouns and three-letter words, but I’ve definitely put in plenty of those. Even though these are all invalid moves, they are movement. This helps me not feel like an idiot just staring at a phone screen. Even better, it can trigger other ideas. It’s a lot easier for me to visualize a full English word when I see part of it, rather than staring at those seven letters in a hexagon. After typing “manu” I found “manual”. There is value in just taking a step forward if you aren’t sure what to do. If you have some kind of goal in mind but aren’t sure how to get there, do something. Take a step towards that goal. If you are trying to get a job in the software industry, especially right now, it can feel overwhelming: so hard, so difficult. But you can take concrete steps towards that goal that are smaller. These steps, when you’re looking at them, may not make 100% sense and won’t make you money. But they are helping you on the path towards employment, and will lead to other steps to take that will get you closer to the goal. I’ve written before about how attending meetups is a fantastic career move . Doing so doesn’t have an immediate payoff, which is super frustrating when you’re trying to find a job, when your bills are piling up and your savings are draining. It’s super hard to think “I’m going to take an hour out of my day and go hang out with geeks talking about ruby “. But joining a meetup can open up other doors. Just like typing “manu” reveals “manual,” going to a meetup repeatedly can help you get a reputation in a community, offer contract opportunities, let you understand the market, and may help you get a job. This is just one example. There are a thousand different steps you can take towards your goal. It can be hard to determine which step is the right step, but the lucky thing is that there are many right steps. You just have to start. Double Down On What Works I like to double down when I have a pattern that is working. In Spelling Bee, if you see a pattern of “ull” being a portion of a word, you should double down on that while you’re thinking about it, and make all of the “ull” words you can. Same with suffixes like “ed” or “ing”. When I see that, I know that I can spell both the initial word like “mint” and get more points by adding “ing” to get “minting”. Even though I have 100% been guilty of a grass-is-greener outlook in my life, where I think, “oh, I’m frustrated with this company, or in this position, or with this situation, it’ll be better over <somewhere else>”. But after 25 years in software development, I know that every place has its issues. I was recently talking to a former colleague, and he was just hired at a super impressive top-tier company. He was talking about some of the issues they have and some of the burning fires. I would not have suspected it from the outside, because they look like they’re killing it. Sometimes, you should look for the good things in your current position and double down on them. This isn’t the same as staying stuck; it’s the opposite of that. Instead, this is about noticing when something is actually working and resisting the urge to abandon it just because it’s familiar. Switching helps when nothing is working; doubling down helps when something is. What does it look like? It might be working a bit extra, taking some free time to study for an employer-funded professional certificate, making small improvements to your team’s workflow, or just bringing your best self to work every day. All of these are taking what works and doing more of it. Double down on what is currently working for you, and it’ll make you happier and you’ll score more points. Three lessons I learned from the Spelling Bee game: switch things up, just try something, and double down on what works. Who knew that looking for words could be so educational?
Sitting on the porch swing and I noticed some friends. Thank you for using RSS. I appreciate you. Email me
One of the most common requests I’ve heard from developers is an in-browser version of Moonshine that can run on a web page. In theory this should be straightforward – we already built MoonshineJS for the previous generation of models, and the core library is written in portable C++, so emscripten can compile it into WASM. There have even been some interesting community porting projects but I held off on official support until I had time to do it justice. I knew that porting the C++ core was just the beginning. Building something that would be straightforward for web developers to use required a lot more: After a lot of work, I finally have a version ready for feedback. The easiest way to try it is on the new moonshine.ai home page, where you can now see everything from a minimal transcription example to a full-blown Granola-style meeting note taker . As an open-source project, all the code for these is available and the simple examples include code snippets in-line too. Here’s one that shows how to run speech to text on a web page, to give you a flavor of the API: You may still be asking yourself why I made supporting Javascript in the browser such a priority? A lot of “X ported to WASM” stories end up being Hacker News bait without having any practical uses. The evidence that drove me was: I’m excited to get feedback on how to improve the initial version, and I’m looking forward to hearing about what people build with it, so please come by our Discord channel if you’d like to join our community. High-level APIs that were both idiomatic for browser Javascript and consistent with the other Moonshine language bindings. Infrastructure for testing from units to full web pages. Integration with the existing CI and deployment process. Examples that were interactive and showed the key capabilities of the library, with interactive inline code. Larger applications that demonstrated and tested how the framework runs in real-world conditions. Improved support for in-memory models and data files. This was involved a lot of changes to the core library, because while there had always been some methods that took memory buffers, coverage was patchy compared to loading from files. Clear developer demand . It came up frequently as a wishlist item when talking to users. Javascript’s dominance . Python rules machine learning, but JS is the most common language for applications, web and server-side. Advantages over alternatives . Voice interfaces are clearly only going to grow in importance over the next few years, but browser APIs are neglected and server-based alternatives are slow and costly compared to our on-client framework. Obvious applications . Dictation and meeting note taking are popular use cases for speech technology already, and talking with AI bots is becoming a lot more common too. Technical alignment . Deep in my bones I know that voice interfaces want to run on the client. The current status quo of streaming audio data to a server just to get text and intent back only exists because models used to be too large to run on consumer hardware. Today even household appliances have enough compute horsepower for local voice agents . It offends my engineering sensibilities to see old approaches linger on purely out of inertia. Speech wants to be free to use and local, just like all our other input devices like keyboards, mice, touchscreens, and cameras. Porting makes that possible on the web. Options . These days a lot of us have to frequently switch between languages and operating systems, and the power of AI coding assistants only increases the pressure to rapidly port applications. A library dependency is a big commitment, and knowing that it will be available anywhere you’re likely to run in the future makes the risk of betting on a framework much lower, even if you don’t need the option in the end.
Seven months of training, 118kg, 2 bad knees, and 2 weeks of notice. Here’s how my first Jiu Jitsu tournament went.
Another month, another agent-shell update. If you missed the last one, have a look at the 0.63 update . While this post showcases the latest highlights, the full list of changes is far chunkier than what we'll cover. agent-shell is a native Emacs mode to interact with AI agents powered by ACP ( Agent Client Protocol ). Since inception in September last year (yikes nearly a year), has featured a shell-like experience, powered by comint mode . There's also viewport mode (via ), if you prefer a more focused experience, and now we have . Chat mode fuses mode with a more traditional chat-like labelling experience. We're living on the edge here, so chat mode is now enabled by default. Ok not really that edgy, it's fairly safe (powered overlays ) and can be disabled entirely via . Chat mode itself is a minor mode, so you can always toggle it on and off via . In the last post, we talked about making less chatty with more grouping for the likes of tool calls and agent thinking, all collapsed by default. If that's far too quiet, you can expand by default via . If you found these two settings either too quiet or too chatty, we now have a third alternative via . When set, only the latest grouped activity is expanded by default, and automatically collapsed when the agent moves on to something else. I've become quite fond of this feature (thank you @nhojb for the PR ), so yes. It's also enabled by default. Queuing received some improvements. The related commands have been consolidated under . You can queue prompts while the agent is busy, then view, resume, or drop pending prompts via , , and . The pending queue is now shown after each new submission. opens a dedicated buffer for crafting a prompt, and it's now more independent of the shell. You can invoke it from any buffer (it resolves to the right shell), sends and returns you to whatever you were doing (fire and forget), and submits the prompt and immediately lets you craft another queued prompt. Shell initialization may take a second or two, depending on what agent you're using, which meant you had to wait for initialization before you could start typing into your new shell. Unnecessary, so that's no longer the case. Shell prompts are now offered as soon as possible, so one can get typing. While Markdown lists are easily digestible without special rendering, we can do better than that, so we now give them a better treatment with normalized padding, indentation, and of course, civilized bullets. TAB navigation made it into fairly early on. I love being able to TAB my way into any section in the buffer and press RET to toggle folding. That's great and all, but we can make the navigation experience richer by welcoming the likes of Markdown source blocks, links, and images to the navigation party. Why these in particular? They are all actionable by RET too, of course. While you'd rightly assume RET opens links to local text files in Emacs and delegates to browsers when needed, Markdown source blocks and images get a similar treatment. While their respective RET actions may not be as obvious, we can certainly make them much more discoverable, so we now add hints. Landing point on an actionable item now echoes a hint of what you can do and which key does it (say, "Press RET to copy" on a source block, "Press + to enlarge" on an image, and so on). Hints are also shown on mouse over events. While offers for customizing image sizes, it's fairly restrictive. Not all images are the same, so why force them all to fit within the same constraint? continues to offer a preferred default, but you can now scale images differently by getting the agent to annotate Markdown images with Pandoc-style link attributes . The attribute block goes right after the image, taking a and/or in pixels or percentages: I may be sweating the small stuff here, but this was really grinding my gears. We have lovely Markdown table rendering, which I'm glad we do as LLMs aren't always great at producing perfectly aligned tables. In the best of cases, the LLMs align the table perfectly, but it's just too wide for our Emacs window. Luckily, our lovely rendering also wraps cells to make them fit into our window. The thing is, all that lovely rendering goes out the door the moment you either resize your Emacs frame or merely split your window, resulting in a monstrosity like this: I know. I'm sweating the small stuff here, but hey we don't have to live like this. Emacs has all the hooks in the world, so let's track window changes and rejoin the civilized world. On a much smaller scale, I also wanted auto resize for images, so here you have it… While we can influence image size at render time, this can still generate undesirable image dimensions, so we can now rescale all images in buffer on demand. Sometimes a hammer really isn't the right tool, so we can now also rescale an image at point. Opening a local file (from a link, image, or mention) now routes through , a standard action you can customize. The default reuses a window already showing the file, or takes over the current one. If you'd like a different window arrangement, you could do something like: Following a local file link now pushes your origin onto xref 's marker stack, so ( ) brings you right back to where you were in your , just like other Emacs jumps. Foldable fragments now use and bind . If you're not a fan of the RET binding to toggle folding, you can now use your preferred binding. If is your jam, you can do something like: If you prefer styling agent thoughts differently, a new face lets you do just that. Streaming performance also received some love. Thanks to @suhail-singh for the profiling and improvements , and to @Scott-Guest and @claytharrison for the trace analysis and benchmarking in #757 . The "Available config options" section now displays possible values. can now be set to a function, letting you compute the available agent configurations dynamically rather than hard-coding a static list. The function is called on every access, so it stays current across code reloads. Maybe you'd like to list only available agents. Here's a rough snippet. now broadcasts an event, handy for external integrations that want to observe streamed output. now works correctly on remote hosts ( #742 by @CeleritasCelery ), smoothing out TRAMP-driven remote agents. agent-shell-hq joins the family, offering an interface for managing multiple sessions. If you peeked at the commit logs , you'll notice I've been working daily on , keeping up with project inflow. Since last month, 27 issues have been closed and 13 pull requests merged. As of today, the backlog sits at 11 open issues and 5 open PRs (versus 13 and 4 last time around). If there's something you'd like me to prioritize, feel free to ping. Vendor-neutral tooling matters more than ever, and there are a couple of ways to help keep going. Some cost money, others just a click. All are appreciated ;) is just me, an indie dev, while the tools it competes with have well-funded teams behind them. Time spent on is time away from work that pays the bills, so if it's useful to you, please consider sponsoring the project. And if your employer benefits from your use, nudge them to chip in too, they can typically contribute at a scale individuals can't. GitHub stars help with exposure, attracting new users and potential sponsors. Starring agent-shell costs nothing and can potentially help bring in more funding, so if you don't mind a couple of clicks, the project can really use another GitHub star . Thank you to all contributors for these improvements! Liking ? Would like to see it evolve? Consider sponsoring the effort. #730 : Queue requests sent during session/push and submit them when it ends ( @catern ) #737 : Add a workaround for goose ( @bergmannf ) #740 : Add a option to ( @nhojb ) #742 : Fix executable-find on remote host ( @CeleritasCelery ) #743 : Preserve "claimed" regions when rendering bold / italic / strikethrough ( @alberti42 ) #746 : Avoid repeated system-sleep load attempts ( @liaowang11 ) #748 : Preserve properties on escaped Markdown punctuation ( @Scott-Guest ) #752 : Guard group member walk against non-advancing block range ( @hamza-m-masood ) #756 : Preserve point during table rendering ( @Lenbok ) #762 : Document viewport workflow ( @KarimAziev ) #763 : Render raw tool output ( @mrychlik ) #765 : Prevent syntax highlighting from delaying other buffers' mode hooks ( @Scott-Guest ) #766 : Refresh viewport header even with nil agent-shell-prefer-viewport-interaction ( @catern )
“The fastest evaluation is the one that never happens.” – Sun Tzu, The Art of Evaluation nixpkgs-multiverse gives you every version of every package that ever shipped in Nixpkgs from a single flake input. Note It continues to blow my mind that this is even possible. It feels like it suddenly unlocks a new dimension of Nixpkgs, and I am still trying to understand what it means. I think this capability is a fundamental change to the way we think about Nixpkgs, and it is not just a new feature. It is a new way of thinking about the entire ecosystem. There was always a penalty at the center of it. Asking for a specific version of , such as , meant fetching the whole ~378 MB Nixpkgs tree from 2021 and evaluating it to determine the . What if we could skip that evaluation? What if we could just ask for the path directly, and have Nix fetch it from the cache if it is there? This is a common idiom if you have ever used . That requires knowing the store path upfront. nixpkgs-multiverse now has a attribute that does exactly that: it gives you the store path for every indexed version of every package. This lets you skip the download and evaluation of Nixpkgs and get the store path straight from the cache. No Nixpkgs is fetched. Nothing is evaluated. No experimental features and no needed for this to work. The complete Nix API, except for releases, works with this fast path. If you want to learn more read the docs about the feature. Every channel bump published a listing of every path Hydra built for it: , or a for back in the pre-2017 era. These files are still available, and they are the source of the multiverse index. The listing is a map from derivation name to store path. The multiverse index is a map from to the revision that shipped it. By joining the two, every historical version gets a concrete address: Knowing the path is not enough, especially in the Nix language. We need to convince Nix that a string that looks like a store path actually is a store path. exists but it is an impure function and requires to work. How do we get around this? We attach “context” to the String. Context is the invisible baggage a String carries in Nix. When you interpolate a derivation into a String, the result remembers where it came from, and that is what makes realise the dependency instead of writing a dangling path into a script. lets you attach it by hand. The identifies that “this String names a store path that must exist,” which is exactly what produces for a path already in your store, except this works for a path that is not in your store yet and is not in this evaluation’s input closure either. Loopole! 👿 We then wrap that in an attrset that resembles like a derivation and the Nix CLI is satisfied: This is tomberek ’s trick from fastpkgs , and it is an amazing trick to circumvent needing . Everything about this remains pure evaluation, and the resulting graph is gauranteed to be bit-for-bit identical to what Nixpkgs would have produced if it had been evaluated. The only difference is that we skip the evaluation of Nixpkgs itself, and instead use the store path directly. The eval path derives the address, the fast path remembers it. A “fake” ( ) derivation has no , because there is no behind it. Nothing can build it and it can only be substituted. The CLI often wants a though when you hand it a derivation attrset, so we must make sure to append the output (i.e. ): and need a real derivation. Every fake derivation carries a lazy that is the real, revision-exact derivation: In the spirit of trying to keep my index small, is empty, so there is not additional information about the package. You can still get the from the real derivation by using as well. This scheme rests on cache.nixos.org still serving thirteen-year-old paths, which thankfully it does and with the same signing key. To demonstrate that the cache is offering nearly every path that Nixpkgs ever built, I ran a census of every indexed version of every package and asked the cache if it was still alive. As of August 14 2026, all 271,187 of them are alive . All of them, down to every NAR payload file. That is 14.8 TB of unpacked software from 2013 onward, one fast command away. 1 There’s some other data on nixmultiverse.com about the census, dependency graphs and additional features. Check it out! All of this is also available via the mvs command line tool as well for offline use. I guess now there is a caveat: there is now a trick in the multiverse. It remains mostly an index, some JSON, and a behind a memo table. The clever trick is tomberek ’s, and it is three important lines. The NixOS infrastructure has never garbage collected the binary cache. It is an S3 bucket that only grows, and the bill is paid by the NixOS Foundation and its sponsors. ↩ The NixOS infrastructure has never garbage collected the binary cache. It is an S3 bucket that only grows, and the bill is paid by the NixOS Foundation and its sponsors. ↩
I’m coming down from spending a few days at Usenix Security, right here in Baltimore. This means that my days have been taken up with two kinds of conversation: first, explaining to colleagues why Baltimore isn’t actually like The Wire. And second: trying not to talk about AI. Here I’m going to break both of those rules. I have many worries about what AI means for our field, for various definitions of “field”. But in this post I want to focus on just one thing I’ve started worrying about, and it’s a perverse thing: specifically, I’m worried that AI is going to make software much too secure. While that doesn’t sound so bad on the surface, there’s a consequence to this. I mean something very specific: I’m concerned that U.S. intelligence and law enforcement agencies are about to go dark, meaning lose a huge portion of their capability. And that this isn’t going to be simply a problem for those agencies, but also for those of us who value computer security and privacy in general. To explain how we got here, we need to talk about recent history. Here we have a real excuse to reference The Wire, which embeds a realistic snapshot of what electronic surveillance looked like in 2002. The cops in that show are after payphones and burners, all used for voice calls. While the mobile phones were new, nothing in here would have surprised a cop from 1989. Less than a decade later, everything was different. The change started in the late 2000s with the rise of smartphones and texting. In 2010, Apple began encrypting iPhone data using a key derived from the user’s passcode, and Google followed behind them. In 2011, Apple deployed end-to-end encrypted text messaging. By 2014, WhatsApp had 600 million users worldwide, and by 2016 nearly a billion — and they were all using end-to-end encrypted messaging. The chart below gives a snapshot of how quickly the world changed between The Wire era and 2016: The FBI and law enforcement agencies noticed the trend and took it very seriously. In 2014, Director Comey announced an initiative called G oing Dark , which would launch a “ national conversation” about what providers could do — or be compelled to do — to make these new communications media legible to law enforcement and counterintelligence. In 2016, the agency stopped talking. When a terrorist attack left the FBI with the shooter’s locked iPhone, the agency ordered Apple to give them access . The company refused . What broke the stalemate — and, to some extent, ended “Going Dark” itself — was something that neither the FBI nor Apple expected. An outside company announced that there was no need for Apple’s assistance: they could simply hack the phone . The Apple v. FBI case turned out to be microcosm of the whole debate. For the next decade, law enforcement and intelligence agencies continued to ask for exceptional access backdoors. But the urgency was gone: agencies and manufacturers knew that law enforcement could purchase targeted hacking tools if they needed them badly enough. Vendors like Apple and Google played a vigorous defense, closing vulnerabilities as soon as they learned about them. But commercial offensive vulnerability hunters consistently managed to keep the edge. Anyway, that’s the history. And now it’s about to be over. In April, Anthropic announced a new model called Mythos that was optimized for software vulnerability finding. The U.S. government temporarily blocked its export, restricting it to U.S. agencies. While the ban was dramatic and made for good PR, it was mostly pointless. OpenAI , along with Chinese open-weight model labs like Z.ai and Moonshot , have since demonstrated that vulnerability finding isn’t anything that a single model can hold a monopoly on. The list of serious vulnerabilities that these models have found is getting scarier (or more impressive) by the day. Initially this might seems like good news for the offense, and for hackers in general. But I doubt it will last. Defenders are now in the process of patching every bug they can find, often with AI helping them. Entire development toolchains are being rebuilt to incorporate powerful vulnerability scanning before software reaches the testing phase. This does not mean that every bug will be found: even calculating the number of bugs in a piece of code is probably uncomputable. In the real world, it does feel likely that we’re going to hit some sort of a ceiling on the number of useful bugs, and probably we’ll hit it soon. Thus: over the next two years, major pieces of software are likely to run out of remotely-exploitable bugs. While I think this is great, for law enforcement and offensive intelligence agencies, it’s going to be a nightmare. For the first time since 2010, law enforcement might experience what it looks like to really “go dark”, across a huge category of advanced (well-maintained) devices and pieces of software. The debate over “exceptional access” mechanisms never really went away. In some places, like the UK, it actually metastasized into something worse. Here in the US it mostly went into hibernation. Some of the slowdown can legitimately be attributed to expert pushback — academics and industry engineers pointing out the risk that backdoors might be abused by the very adversaries they’re designed to protect against. But I fear that the market was just pricing supply. The destruction of the low-hanging vulnerability fruit will make law enforcement (and intelligence) agencies’ need much more acute. The demand for constructed, intentional backdoors will begin in earnest. There will be enormous pressure on industry to re-architect their systems to make their systems friendly to exceptional access. In some cases, governments will ask for these capabilities in the expectation that they’ll be useful for spying on other governments — a strategy that might have been undetectable in the pre-AI era, but that probably will be detectable now. This might result in other governments curtailing their dependence on US software. In fact, the worst part about this dynamic is that these potential new backdoors will begin primarily useful for allowing the US to weaken its own systems, which will in turn allow foreign adversaries to find new ways to attack our communications. This deliberate self-sabotage will happen just at a moment when we’re finally learning how to defend our own infrastructure. I honestly have no idea. This is not a call to action for experts to rally behind a sophisticated plan. Like so many things about the AI revolution, it’s just occurring to me that we’re on a long greasy slide to a place that will look different than where we are today. Just realizing this doesn’t mean that I have a clever plan to avoid it. In this case, we’re just going to have to hope that this time we make the right choices, for no other reason than that they’re right.