Latest Posts (20 found)
Unsung Today

“This way, the interface does not get in the way.”

Ilya Birman on his blog talks about interfaces that unnecessarily slow people down, in a series of two posts. In the first one , titled “Let me click,” Birman shows a few places that force the user to go through a roundabout series of clicks, instead of taking a direct route: = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/this-way-the-interface-does-not-get-in-the-way/1.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/this-way-the-interface-does-not-get-in-the-way/1.1600w.avif" type="image/avif"> In Aegea’s comment settings, there is a “send by email” checkbox with an email-address field associated with it. If the checkbox is unchecked, there is no point filling in the field: the address is not needed for anything else. If the checkbox is checked while the address is blank, the system cannot send anything as it does not know the address. In short, these controls are interconnected. Logically you could disable the input altogether if the checkbox is unchecked — there is no point filling it in anyway. But that is irritating. What if I want to enter the address and then turn on the checkbox? It would be even worse not to let me turn off the checkbox when the address is filled in. I want to turn it off — let me click! In a follow-up , Birman talks about a specific example from the podcast app Overcast, whose creators faced with a tricky systemic challenge: For any podcast, you can set how many episodes to keep downloaded on the device. Say you set the limit to three, then stop listening to a podcast regularly: the next three episodes download, and after that it stops downloading them — why waste the space? […] [But,] what if someone manually asks to download an episode when they already have three downloaded? You can’t just immediately delete it to keep things tidy. […] Finally, Marco’s wife Tiff suggests a solution to all the problems: just don’t let users download more episodes, she said, and show them a message, roughly: “Your episode limit is three, but this would be the fourth, denied”. Marco liked the solution. I didn’t. Sure, it solves Marco’s problems, but not the user’s. I haven’t seen the Overcast block in action, but all of Birman’s examples across both blog posts rang true to me. Motor memory wants what it wants, and stops for no one. I wanted to add two things from my end. Birman presents this as “Let me click,” but I wanted to offer two alternative/​complementary principles that helped me before: I agree with all the examples given by Birman, but things can be stranger, and sometimes letting people click or do things in any order can make the interface harder to understand. In Figma, each text box dimensions could be set to be completely automatic (automatic width and height – used for short labels), with only specific width (and automatic height – used for paragraphs of text), or with manual width and height (used for graphic elements): You can also see that in text boxes where width or height are automatic, the fields for those values are grayed out, and cannot be changed. In order to change them, you have to switch to a particular mode first, which makes the relevant fields active. This is to help people understand how those things relate to one another in a pretty tricky space, but it effectively means sometimes Figma won’t let you click, and will force you to do this in one specific order. But this is only in those fields – you can always grab the object on the canvas and resize it, and it will switch to manual on its own, respecting the drag’s momentum: The inconsistency here is intentional. I’m not saying these are the right choices, and we indeed heard from some users frustrated that they cannot click easily when they already understand the system (unfortunately, I am aware of no good affordance for “breaking” a disabled field in modern GUIs, like a molly guard ). I mostly wanted to share that these things can be hard, and the balance between “let me click” and “not being able to click can be helpful in understanding the system” tricky to achieve. I keep thinking of the story from a decade ago when someone’s phone rang at the front row of a New York Philharmonic’s concert , prompting an actual performance halt, and anger from both the conductor and the audience. The ashamed patron’s eventual explanation was: “I turned the phone to silent, but I also had an alarm set up.” Let’s assume this is actually true (in people’s reports, the phone rang the marimba ringtone, which isn’t standard for alarm – but people’s recollections are routinely flaky, too). This, I think, is a perfect example of “let me click” in action. You can set an alarm for a certain time, and you can subsequently turn the phone to silent. You can also do it in the opposite order. You just gave your phone two inconsistent instructions, and the phone logic decided in either case the alarm will win. I can’t think of an easy interface solution here. It feels correct for the order to not matter here. It’s hard to imagine disallowing you switching to silent mode with any alarms on, since the silent mode also affects calls. It’s also hard for me to imagine any effective UI warnings at any given moment in the process. (Besides, back in the day, iPhone used to give you a gentle one via a little alarm icon visible in the top bar.) = 3x)" srcset="https://unsung.aresluna.org/_media/this-way-the-interface-does-not-get-in-the-way/5-framed.1600w.avif" type="image/avif"> The silent mode and the alarms are, as Birman put it, interconnected – but their connection is ultimately tricky to explain to the user. #case study #details #flow #system design Let me do things in any order . Here’s Google Home app, where I can change the temperature and the time for holding – but if I change the time first, it frustratingly resets when I subsequently change the temperature. It forces one specific order in an interface that suggests any order is okay: The current action has the most momentum . The user is right there, active, tapping on things, wanting to get stuff done. If Overcast indeed throws a “your episode limit is three” message, then it forces the user to remember how to get to the settings and change it. The momentum is lost. The decisions of past me should not be as important as the decisions of present me.

0 views

Fragments: August 24

I was listening to Ezra Klein’s interview with Helen Toner about the recent OpenAI hack of Hugging Face and the subsequent discovery that there were swarms of agents inside OpenAI doing unsanctioned activities. One of the points Klein made was that at no point did any of these (thousands of?) agents ever try to check in with a human [Klein:] So these message boards — you have however many A.I. agents posting hundreds of thousands of messages. At no point do they say: Hey, researchers, programmers, parents at OpenAI, Anthropic — do you want us coordinating with each other on this message board we have created in the innards of your systems? [Toner:] Or even F.Y.I., we have a message board we’re coordinating on in the innards of your system. Listening to that, another thing occurred to me - none of these agents thought to rat the others out . No “hey, some of the agents in here are doing sketchy things”, no sign of an AI whistleblower. ❄                ❄                ❄                ❄                ❄ Is the AI bubble so big that the frontier companies like OpenAI and Anthropic have no way of becoming a viable business? If that’s the case, Bruce Schneier and Nathan Sanders have a possible path: Evidence suggests the market itself could reassess that these companies offer nothing of financial value. In that case, perhaps we can return them both to their original purposes. If these AI companies should fail in the financial markets, the US should nationalize them and convert them into national labs operated under democratic control that preserve their benefit to the public interest. Such an idea may strike many people, used to the laissez-faire free enterprise world of Silicon Valley, as sacrilege, disaster, even socialism. But the United States made world-beating technological progress through such institutions in the recent past. AT&T was a quasi-government entity that led the world in telecommunications and electronics after the second world war. The US has a long, successful history of these kinds of institutions, which have produced world-shaping innovations in spaceflight, telecommunications, nuclear power and more. Congress currently manages a $200bn R&D portfolio, within which frontier AI development is, arguably, a glaring gap. ❄                ❄                ❄                ❄                ❄ Here’s a message for those readers who live in Massachusetts, just to the north of me, specifically in congressional district MA-06. I don’t usually endorse political candidates, but I’ve made an exception for Beth Anders-Beck , who is running for that house district. I’ve known Beth for many years and have a high opinion of her smarts, wisdom, and compassion. They would make an excellent member of congress. ❄                ❄                ❄                ❄                ❄ Kevlin Henney posts “one weird trick” for deciding when to skip reading LinkedIn posts , essentially by identifying a common pattern for skippable posts: It seems like a good approach. I, however, have a simpler one - skip all LinkedIn posts. ❄                ❄                ❄                ❄                ❄ Bartosz Ocytko has detailed and thoughtful post about the usage of agentic programming at Zalando . Like most companies I hear from, they are convinced of the value of agentic programming but still exploring how best to do it. One notable step they’ve taken is building platforms to act as a clear portal for API access and tools to support chat UI and CLI. This allows them better support good security practices and to monitor usage of models. They have seen signs of agentic programming increasing the complexity of codebases, including leading to larger commit messages. The write-up spends a lot of time on knowledge sharing, how to pass on skills, and the support of experiments. With >200 teams innovating and broadly exploring the ecosystem, the question arises whether and when to converge. We believe it’s way too early for this. While agentic engineering practices are still in their early stages, our key objective is transparency and exchange across teams. I was struck by their use of an LLM to assess the risk of pull-requests. Those with a low risk of rollout can be auto-approved, reducing lead time by 20-40%. An interesting consequence of this is that it encouraged folks to split pull-requests so low risk portions can take advantage of the fast approval. Any changes to configurations are automatically made high-risk, which they feel protects them from common outage traps. They repeat the common thread that the value of AI depends greatly on underlying skills. Like anyone in the industry we observe how AI amplifies the good and bad practices across our organization. Teams that get carried away with agentic engineering end up with large PRs that discourage reviewers and slow down delivery until a team adjusts their practices. ❄                ❄                ❄                ❄                ❄ Julia Curlee was a senior intelligence official in the White House. She had served under administrations of both parties, been the briefer for Vice President Pence, and on the National Security Council under Biden. She writes an absorbing account of her relationship with Pence and shares observations about the changes to the intelligence community under the current administration, including recent events at the CIA (gift link) The agency has been gutted as part of a deliberate plan, the director of the Office of Management and Budget once boasted, to put the people who defend our country “in trauma.” Analysts have been fired in public or questioned by the FBI; decade-old assessments have been denounced by the CIA director in the press. The president calls analysis “virtual treason” when it contradicts his preferred reality, and uses the CIA to undermine public confidence in American elections. Fear has done its work. Irreplaceable officers with crucial language and technical skills, and decades of experience, have walked out the door. Those who remain within an agency built to deliver hard truths are being muzzled. For a worthwhile sample of her analysis, read this evaluation of the current bargaining between the US and Iran Most wars do not end in “unconditional surrender.” They end when both sides accept terms. Paul Pillar’s classic study of war termination, “Negotiating Peace,” treats combat and diplomacy as a single process: Each side fights to improve the terms it can demand at the table, and talks to lock in what the fighting has won. She continued to serve the second Trump administration even though they knew she was trans, until her position was made public. Autocrats seem appealing, with the promise to get things done without the ponderous constraints of rule of law or bureaucratic procedure. There are occasional “Good Emperors” who raise people based on merit, but more often such power attracts corruption, nepotism, and toadies. Flailing regimes dehumanize minorities to distract from their failures. When the economy collapses or a war goes badly, they find a tiny group of people, make them the enemy within, and rally the country against them. This is how it’s gone in Iran. Hungary. Russia. I wrote PDBs about it. This will not stop with trans people. It never has. Post is too long Contains a (crummy) info-graphic No voice of poster (instead “aspiring anodyne anonymity”

0 views

The Red Mailbox

The Red Mailbox next to our village’s primary school is my primary drop-off point for writing letters . Over the decades, its bright deep red brilliance has been gradually replaced by a patina of sunburned broken red. And yet, The Red Mailbox persists. It still exists. It existed over thirty years ago, when I went to that very school next to it. Other Red Mailboxes aren’t that lucky: in 2019, BPost—the Belgian posting company that was privatised in 2000—removed over one fourth of the Red Mailboxes all over Belgium in an attempt to “save the company”. The biggest reason might not surprise you: most of these boxes didn’t receive much letters: Volgens BPost is het aantal brieven dat mensen in de rode brievenbussen deponeren, met 60 procent gedaald sinds 2004. Uit een kwart van de bussen haalt de postbode nog hoogstens zes brieven per dag op, luidt het. [According to BPost, the amount of letters that people deposit in the red mailboxes has diminished by 60 percent since 2004] The Red Mailbox. Emptied at 10:00 AM. That year also happened to be the year of BPost’s stock market crash. Still, I think ultimately their decision was the right one: of all the letters I send out, I only receive about a fourth replies in that same analogue form. Especially in 2026, people don’t use The Red Mailbox anymore, turning the battered broken red metal box attached to a brick wall into a weird artefact of the past. I wonder how long it’ll take before the same emotion is triggered as wandering around in The Legend of Zelda: Breath of The Wild ’s broken world, where artefact bits and pieces poking out of a ruin reminiscent of a once thriving community now only cause weariness. Before the Mailbox Removal Program, BPost also permanently closed multiple post offices. Our village’s old post office building now is a Turkish food joint. Yet another thing of the past, just like physical banks, toy stores , and if we’re not careful, bakeries and butcher shops. Yet we are nimble: a bike ride in two directions, both about four kilometres, will take you to another office. Do verify opening hours before leaving. Even Google Street View archives dating back to 2009 do not have a photo of the old post office building in its original state. The archives do contain traces of deceased Red Mailboxes, such as the one below weirdly enough mounted against the front facade of someone’s home? I would like to believe that the owners of the house kept a notebook besides one of the windows to track all those weird people stopping by to drop their weird letters. An archived copy from Google Stret View of the one that disappeared, mounted next to a red drainpipe. The dismounting job revealed an ugly square stain behind it. Passengers will no doubt wonder what used to be on that wall, thinking the owners of that house must have done something weird. They didn’t: the government did. I wonder what personal history The Red Mailbox can tell us. How many love letters did it ingest all these years? Or birth announcement cards? Marriages? The bearer of happy news. Let us intentionally leave out the boring and perhaps more depressing tax related correspondence. We made good use of that very same mailbox when we got married, and when we had our daughter and son—I can distinctly remember it barely containing the envelopes when I dumped all those birth cards in there at once. The expected Thunk! sound of the envelope hitting the bottom was replaced by the shuffling of papers and me trying to jam them all in. I still make good use of that trusty old mailbox whenever I feel like using a pen to send a message to a pen pal, which admittedly happens less and less. Perhaps I too am a part of the problem. The Red Mailbox was not always there. In fact, it’s one of the more modern alternatives—relatively speaking—compared to the late 19th century cast iron standalone models that are now officially classified as being part of our Flemish/national heritage . This particular model in the link still stands proudly in Antwerp, yet these are the exception to the rule: I haven’t encountered a cast iron painted one in our vicinity. The little door to retrieve the letters is a lovely touch. According to heritage info of the city Spa , these models were cast by the J.G.Requilled foundries of Liège between 1860 and 1938. The most striking difference between these original models and the ones on the photos above is perhaps the fact that its purpose is no longer explicitly mentioned: the French inscription “LETTRES/IMPRIMES” is gone. By now everyone knows that a Red Mailbox is a letterbox from BPost, not a personal one. I hope. If you ever received a letter from me: think of The Red Mailbox where its humble journey all the way to your letterbox started. Related topics: / letter writing / By Wouter Groeneveld on 24 August 2026.  Reply via email .

0 views
Unsung Today

“This would result in some serious street cred and respect.”

A fascinating 11-minute video by Modern Vintage Gamer about the PC videogame piracy in the 1990s: = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/this-would-result-in-some-serious-street-cred-and-respect/yt1-play.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/this-would-result-in-some-serious-street-cred-and-respect/yt1-play.1600w.avif" type="image/avif"> This is a “no honor among thieves” story about a few groups racing to pirate and release one of the most important videogames in history: 1996’s Quake . There is a lot of interesting vernacular here: leetspeak , nfo files , cracktros , ASCII art, BBSes, IRC, a bunch of slang, and even specific font styles. At the same time there is also some unexpected bureaucracy – apparently videogame pirates had sort of a governing body, which perhaps was also somewhat corrupt? #games #piracy #youtube

0 views
Unsung Yesterday

Testing tip: Make your keyboard fast

I believe every modern operating system allows you to set the key repeat rate and the delay before first repeat: = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/testing-tip-make-your-keyboard-fast/1.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/testing-tip-make-your-keyboard-fast/1.1600w.avif" type="image/avif"> = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/testing-tip-make-your-keyboard-fast/2.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/testing-tip-make-your-keyboard-fast/2.1600w.avif" type="image/avif"> By default, on a Mac, the initial delay is 500ms, and the key repeat rate 1000ms. You can adjust both: On a PC, the shortest values are exactly the same, but Windows 11 only allows up to 405ms between key presses, and 1000ms of initial delay. My suggestion is to make them both as short as possible for testing. Why would that be helpful? There are two reasons. First, it’s good to test whether your keystrokes behave well when repeated. Occasionally you might want your interface to suppress repeating, or do something special if the key is held longer. Second, it’s a good test not just of repeating, but just the user pressing keys really fast. It’s very important for the UI to never make you wait for any animations or transition, and a lightning fast 30fps repeat rate is a good way to stress test your system this way. Here’s me holding the Tab key in Figma at a standard key repeat speed: It’s nice to see those transitions help you orient yourself as you’re thrown around the canvas. But look at what happens when I do the exact same thing with the key repeat cranked up: The transitions are now slowing the UI so much that you can no longer see any canvas movement – the canvas only catches up when I release the Tab key. The smooth movement assumed a certain minimum key rate, and perhaps wasn’t tested with a faster one. There are of course ways about it, like speeding up the interactions, suppressing the transition smartly, or introducing a custom repeat rate if needed – I talked about it a bit in my essay about designing fast keyboard interfaces . Here’s an example of an interface that suspends transitions at a fast keyboard rate and responds in real time: This looks very chaotic especially since you’re not in a driver seat, but this interface is at least honest and doesn’t make the user wait. And it might not be pure madness – you would be surprised how fast people’s brains and fingers can react. But first, you have to be aware of the problem, and I think setting up those key repeat values as short/​fast as possible will help you find more of those kinds of issues. (And you might choose to keep those high speeds anyway in regular use, like I do.) #flow #keyboard #tips Initial delay goes from 250ms, to as long as 2s. Repeat rate goes from 2000ms (once every two seconds) all the way to a whopping 33ms (30 times a second).

0 views
Unsung Yesterday

“A quick internet search should provide you with step-by-step instructions.”

In 2018, Tom Cruise and Christopher McQuarrie teamed up to make a short PSA about the dangers of frame rate interpolation, a.k.a. the soap opera effect : = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/a-quick-internet-search-should-provide-you-with-step-by-step-instructions/yt1-play.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/a-quick-internet-search-should-provide-you-with-step-by-step-instructions/yt1-play.1600w.avif" type="image/avif"> It’s a strangely boring video from the men known for exciting filmmaking, but it captures a fascinating debate. TL; DR of their argument: Home TVs have an option to take a typical 24fps movie and create interim frames to make it feel like a 60fps production. It’s often on by default, but you should turn it off. What’s interesting is that 60 or more fps is indeed in some ways objectively better: you see smoother movement, and you can notice more. Once you start paying attention, any whip pan or drastic movement in the cinematic 24fps appears very choppy… …and, of course, that is completely missing the point. The argument for 24fps is that it’s simply cinema’s vernacular, right next to anamorphic lenses with their blue lens flares and vertical stretching. The movies are not about conveying information, but they are about conveying a certain feel . (And also, about tradition.) I’m mentioning this because I spotted a similar battle happening when it comes to scrolling. Go to any page in Vivaldi and hold a down arrow key for a while. That scrolls through the page, moving it in chunks, reacting immediately to the initial key press and then the synthesized autorepeat presses: [force height limit] Firefox listens to the keys the same way, but it “upgrades” the choppiness to a 60fps movement by interpolating it: [force height limit] And Safari does something different altogether – it approaches it more like a game physics engine would. It only listens to ↓ down (when it turns on the scrolling “motor”) and then ↓ up (when it turns it off): [force height limit] This has an interesting effect: the movement is smoother than even Firefox’s – that’s because it doesn’t try to straighten something choppy, but it is itself smooth, by nature. At the same time, this approach feels a bit loose, like driving an old 1960s car. It also takes away any control of speed; Safari pages always scroll at this rate, no matter your keyboard settings. (Arguably, however, control via the key repeat rate that other browsers respect is also an illusion, as it affects typing as well. Would you ever change it just to control the scroll speed?) Of course, the analogy to frame interpolation doesn’t really make sense. Safari is not a soap opera, Firefox is not a smooth motion effect on your TV, and Vivaldi is not a cinematic experience. In contrast with movies and television, I don’t think there are any expectations or tradition here. In this particular context, I believe Firefox and particularly Safari are better because this is about conveying information, and smoother scrolling does help your eyes and your brain connect all the little befores with all the little afters. But I am sharing it mostly as a reminder that a keyboard is really just a button board by a different name. And sometimes it’s good to look at a key hold, and decide: #interface design #keyboard #motion design #youtube is this a sequence of pulsating key presses, or is it a button being held down and then released up?

0 views
Kev Quirk Yesterday

Dungeon Crawler Carl

Author: Matt Dinniman Genre: Fantast, Sci-fi Released: 2020 Rating: ★★★★★ You know what’s worse than breaking up with your girlfriend? Being stuck with her prize-winning show cat. And you know what’s worse than that? An alien invasion, the destruction of all man-made structures on Earth, and the systematic exploitation of all the survivors for a sadistic intergalactic game show. That’s what. Join Coast Guard vet Carl and his ex-girlfriend’s cat, Princess Donut, as they try to survive the end of the world—or just get to the next level—in a video game–like, trap-filled fantasy dungeon. A dungeon that’s actually the set of a reality television show with countless viewers across the galaxy. Exploding goblins. Magical potions. Deadly, drug-dealing llamas. This ain’t your ordinary game show. Welcome, Crawler. Welcome to the Dungeon. Survival is optional. Keeping the viewers entertained is not. Learn more on Goodreads ➡ Started reading this book after a couple of friends in work recommended it to me. Couldn't put it down and read it in a couple days. I've gone straight onto the next book in the series, and I'm really looking forward to seeing where it goes. The book has Hitchhiker's Guide To The Galaxy vibes in that it doesn't take itself too seriously, and is a lot of fun to read. Highly recommended. Thanks for reading this post via RSS. RSS is ace, and so are you. ❤️ You can reply to this post by email , or leave a comment .

0 views

How To Report A Bug So It Actually Gets Fixed

I wanted to make a blog post like this for a long time, because it’s something that I wish I could find more of myself. I think one of the things that helps us most in our careers as software engineers is knowing how to debug problems, how to reproduce them, and how to report them. What prompted me to write this was watching this awesome video from Kovarex, the founder of Factorio, where he goes through a bug report and tries to fix it. I thought the bug report was written pretty well, and I thought it might be helpful to show how I went about writing a bug report like this myself, and what the thought process was.

0 views
Pete Warden Yesterday

How to see hidden East Anglia as a US service member

River Lark and old mill buildings, Mildenhall  by Bikeboy Today I saw a question posted on Reddit from someone who’s deploying to Mildenhall from the US. They wanted insights on how to dive deeper into English culture as a visitor, and I since I grew up nearby and I’m waiting on training I ended up writing a long answer . I’ve no idea if search engines even work any more, but in the spirit of leaving breadcrumbs for anyone with similar questions, I wanted to share a version here too. As someone with family still in Mildenhall and who grew up nearby, but who has lived in the US for 25 years, here are some thoughts for whatever they’re worth. The area has hosted Americans for over 80 years, so the locals are very used to their presence. I even had some USAF kids in my school classes, from families who wanted them to mix outside of the base. I still remember the day one brought in his dad’s pressure suit from a Blackbird for our equivalent of show and tell. I don’t know whether going to an English school was something they benefited from, but I believe it’s still an option if you’re interested. For us it was great to have friends who could bring us American candy. The house rental market is very much set up for visiting military, for better or worse. You’ll have no problem renting, the landlords know they can get help from the base staff if there are any issues, but you’ll also find most places are set up for fairly short term deployments of a couple of years or so. This means you’re not likely to get a rose-covered cottage, but something more functional. You’ll also find it hard to get away from other Americans without traveling a few miles. Photo of the Norfolk Broads by Russell Smith Don’t let any preconceptions of the English countryside mislead you. I find the Fens beautiful, but they’re basically drained swamps, very flat and agricultural. I highly recommend experiencing it from the water, with a narrowboat tour through the Norfolk Broads. You can actually travel pretty much the entire country by connected canals, and some people live on boats year-round. It’s a blue-collar rural area with all the same drugs, violence and bigotry you find in similar areas in the US, though maybe not as obvious. I’ve lost too many friends to suicide and drunk driving, and ask a local about Roma people if you want to hear some ugly attitudes. Different pubs are focused on locals, tourists, and service members. You’ll probably find the locals’ pubs give you a bit of a cold shoulder at first, but that’s also where you’ll get the most authentic interactions if you stick it out. You should also check out the local churches. Even if you’re not religious, you’ll be welcome at a service, and you’ll find them amazing places to visit. I’m an atheist, but Over (confusing village name) St Mary’s is still a special place for me, with a documented history back into the 1000’s, and parts that likely predate that. It’s also been described as having the best gargoyles west of Notre Dame! Churches generally remain unlocked during the day, and if not, the graveyards can be beautiful and fascinating in themselves. There are often local history guides available for purchase (unattended on the honor system) inside the churches. On a similar note, search for local archaeological sites, there are often visitor days or even chances to volunteer. Maybe you’ll find a bog body or a hoard of treasure ? West Front, Abbey Gardens – Bury St Edmunds, by Jim Linwood There are also a lot of small towns with amazing history and architecture, like Bury St Edmunds and St Ives, that don’t have the crush of tourists you’ll find in Cambridge. Look up market days too, it’s often a lot of tat but mixed in there can be some gems. They tend to be aimed at bargain shoppers too, not bougie like farmers markets here. St Ives’ market is over 900 years old , and still held weekly on Mondays, with special events on Bank Holidays. There’s a weird love/hate relationship with the US. I had friends who insisted Americans aren’t funny, while being addicted to shows like the Simpsons (I’m dating myself I know). There’s also strange holdovers like believing there’s still a “special relationship”, the US is “over the pond”, and Americans are our “cousins”. You’ll likely get a lot of questions (or digs thinly disguised as questions) based on stereotypes, news headlines, and people who think their holidays in Orlando give them in-depth knowledge of the US. There’s still a lot of genuine affection for America though. World War II looms large in the British psyche, with war memorials listing the dead in most villages, and we know it would have gone very differently without American help. Bury St Edmund’s War Memorial I wouldn’t worry too much about American habits, but it can be hard to gauge British levels of friendship as an outsider. You might hear “we should get a beer sometime” and try to set a date, only to discover that was just a polite way of saying you’re on good terms with them, like “we should get lunch”. I’ve also had British friends complain that Americans behaved in a way that was perceived as being very close, but then failed to stay in touch with Christmas cards once they left. I feel like the US is a lot more set up for people moving a fair bit, and Britain has a lot more people who live where they were born, though part of that may be because I’ve always lived in big cities here. On the surface a lot will feel familiar, especially because of the shared language, but with a bit of patience and exploration off the beaten track, you will start to see how foreign and complex the UK actually is. I think you’re approaching this with a great attitude and hope you and your family have a good time.

0 views
Unsung 2 days ago

YouTube’s right click relay

An interesting mechanism in YouTube I just learned of. If you right click once, you get their context menu: = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/youtubes-right-click-relay/1.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/youtubes-right-click-relay/1.1600w.avif" type="image/avif"> If you right click again , you get the native browser menu: = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/youtubes-right-click-relay/2.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/youtubes-right-click-relay/2.1600w.avif" type="image/avif"> This is clunky and not discoverable, but I can see YouTube team’s bind here. As far as I know, it is not possible to extend browser’s native right click menu (even with user’s consent), or invoke it in some other way (so that, for example, they could have an entry point to the native menu in their menu). At the same time it does feel like the correct use of a right click menu, to host quick functions like Miniplayer or Copy Video URL At Current Time or Copy Embed Code – and, you can also see how Chrome’s own menu has a lot of cruft in it. In a way, this might be what is often called a “progressive enhancement.” It is better than blocking the native menu altogether, and I cannot think of any smarter alternative given the constraints. But it is clunky, as things often are when websites venture out to become web apps. #mouse #web

0 views

Concurrent Servers: Part 8 - Go

This is part 8 in a series of posts on writing concurrent network servers. In this part, we'll switch to Go and see how it tackles the challenges described earlier in the series. All posts in the series: This post assumes a basic familiarity with the Go programming language. As before, we'll start with a sequential server for the basic state machine protocol presented in part 1 . This is the main function: As in the previous parts, the server is "infinite"; it never stops serving new connections until it's explicitly killed. This is the function implementing the protocol for a single client; it takes a net.Conn value that represents a socket with a client connected on the other end: Rather than directly exposing OS threads, the Go runtime implements its own M:N scheduling of lightweight goroutines on top of OS threads. Using goroutines in Go is cheap - both in terms of syntax and developer effort, and in terms of system resources . Here's a version of our serial protocol server that serves clients concurrently by launching a goroutine for each client. The part of the code that's different from the previous sample is highlighted: The concurrent modification in this case is particularly simple because the server is infinite; there's no point waiting for these goroutines to finish (and hence no need for a sync.WaitGroup ). The parameters for server.ServeSerialProtocol are lexically captured from the enclosing scope and its return value is handled by the surrounding closure. Because goroutines are very cheap, this server is very unlikely to run out of resources due to launching too many goroutines; in fact, it will probably run out of something else - like file descriptors for sockets - first. However, sometimes it's still useful to limit the degree of concurrency - even in Go, and we'll discuss some approaches to do so in the following sections. Here are some scenarios in which it makes sense to limit the degree of concurrency in Go programs, even though goroutines are cheap to launch and operate: Let's switch to the primality testing server from part 4 for the rest of the post, because it represents a somewhat more realistic workload. As a reminder: the server receives numbers, simulates blocking by sleeping, and returns "prime" or "composite". The unbounded one-goroutine-per-client version looks almost identical to the previous code sample, except that the goroutine invocation calls another function: Where ServePrimeProtocol is [1] : The simplest way to limit concurrency in Go is by using a counting semaphore, implemented with a channel: The channel sem serves as a semaphore; note that it's a bounded channel with a maximal size. A token is acquired by sending to the channel, and released by receiving from the channel. When the channel is full, the send operation sem <- struct{}{} blocks until a token was removed by some other goroutine [2] . The type of the channel is struct{} which means "empty", or "no data". This is idiomatic in Go for channels that are used solely for their semantics, not to send/receive any actual data. Since launching goroutines is cheap and limiting concurrency is easy as shown above, the "worker pool" pattern is often unnecessary for scenarios like our server. Still, it has occasional uses (such as when workers have to maintain some non-trivial state across tasks) so it's worth discussing it here. Here's a variant of our primality testing server that uses a worker pool: A fixed number of worker goroutines is launched; these goroutines all receive "jobs" from the same channel. In the Accept loop, each client connection is sent to this channel as a new job and is picked up by the next available worker. As mentioned before, you would typically see a sync.WaitGroup somewhere to ensure clean shutdown of goroutines, but in our case it isn't necessary because we have a server that never exits. Do programmers have to resort to async / event-driven programming in Go? In my experience, almost never. Go was designed from the bottom up to be suitable for large-scale concurrency; goroutines are very cheap to create, have a tiny memory footprint and switching happens very quickly, all in user space. Measurements I ran back in 2018 have shown switching times of ~170 ns, as compared to 1-2 microseconds for threads on Linux. Moreover, Go already uses event-driven loops like epoll underneath for I/O. Goroutines that wait for I/O like sockets are effectively "parked" and consume no resources (beyond their small memory footprint); they are woken up by Go's runtime when their I/O descriptors are ready - this is very similar to how asynchronous programming works! That said, some people certainly do try to stretch their resources even more with direct asynchronous programming in Go when millions of streams are handled concurrently. All I'll say is that this is very rare, and an overwhelming majority of users never have to do this. In 2018, I wrote a post named Go hits the concurrency nail right on the head , and after several more years of active coding, I fully stand behind that statement. Go is extremely powerful and ergonomic for concurrent programs; while other environments go to great lengths to implement async-await style event loops in libraries, in Go it's already baked into the core language and runtime. You want event-driven I/O with very lightweight green threads that can also execute blocking tasks without worrying about the function coloring problem ? Go has you covered. All the code for this post is available on GitHub . Careful readers will note two issues with this code: (1) the protocol assumes the complete number is read from the socket in a single conn.Read call, and there's no framing - separation between distinct numbers; (2) the prime checking loop uses i*i which may overflow for large numbers. These issues are consistent across all the versions of the prime server in C, Python, JavaScript and Rust in earlier parts, because my focus was on the simplest possible code to demonstrate a point about concurrency. Part 1 - Introduction Part 2 - Threads Part 3 - Event-driven Part 4 - libuv Part 5 - Redis case study Part 6 - Callbacks, Promises and async/await Part 7 - Rust Part 8 - Go (this part) Tasks may be compute intensive, and the CPU capacity of any server is inherently limited. If too many concurrent goroutines compete for limited CPUs, they will all make very little progress. It may make more sense to have fewer tasks that complete in a reasonable time. Protecting potentially limited downstream resources, such as concurrent DB connections or other services. For example, if the server has to send requests to other services for each task, and these are rate-limited, concurrency will have to be carefully managed. Security reasons when work is dictated by clients; malicious clients can overload and crash a service that's too eager to serve, making it unavailable for legitimate clients.

0 views

A Syncthing and SQLite Gotcha

So, I have this little app, Epoch , that I use to keep a journal. It’s a tiny Rust web app that runs as a systemd service and uses SQLite as the database. I use a desktop and a laptop regularly, and use Syncthing to synchronize them, including Epoch’s database. That way I can use the app on both devices without needing a server to synchronize them, the tradeoff being that I have to make sure the sync is finished before performing any mutations. But I had this bug. Say I edit today’s entry on the laptop, come home, wait for Syncthing to finish, then I’d open today’s entry on the desktop, and the text would be missing. It’s not that the server is holding a lock on the file and preventing the sync: opening the database with the command line tool shows the new text is there. If I restart the server, Epoch can read the new text. My mental model was: The rusqlite object points to the database file. Syncthing swaps the file’s contents from under it. Subsequent queries go to the new file. Turns out there’s a very important part of POSIX filesystem semantics I was ignorant of. The standard way to replace a file safely (i.e. atomically) is the system call: Which Syncthing uses. This I know. What I didn’t know is: what happens if other processes had open file descriptors pointing to ? Do they see the new contents? No: those processes can keep reading and writing to the old file object , but the file is orphaned in that no path points to it. And once all file descriptors are released, the file becomes inaccessible. I’m used to thinking of filesystem operations in terms of “this syscall takes a path and gives you a pointer to the file, which you mutate directly”. Whereas works at the level of directory entries: it atomically mutates the mapping from pathnames to files but doesn’t touch files at all.

0 views
Ahead of AI 2 days ago

How Claude Watermarks AI-Generated Text

I recently posted a Substack note about Claude’s new watermarking process and implementation. Since it’s such a popular topic and sparked such a lively discussion, I thought it might be interesting to go into a bit more detail when explaining how it works. Instead of the usual text article, I recorded a little lecture on the topic (to change it up a bit from my usual articles). So, below is the video along with a transcript. Originally, I planned to make 10 slides and record a short 10-min video. However, while putting it together, I added some crucial details here and there, resulting in >50 slides and a 48 min recording. I hope that this now explains it well, though! Happy watching! I also have a YouTube version if you prefer using the YouTube player And here is a link to the slides Subscribe now Note: The transcript below is slightly edited and cleaned up for readability but preserves the overall order and flow of the video lecture above. Slide 2 of 52, time stamp 0:00 Hi everyone. So, a few days ago, Anthropic announced that they will watermark the text outputs of their Claude models. I then did a social media post briefly explaining how that works. And yeah, this was quite the popular post. So not the watermarking itself was popular, but I guess the explanation or the mechanism behind it. Then, it might be worthwhile expanding this a bit to explain it in more detail, because this post only had one figure, and there were a lot of questions and discussions. So, I thought, well, let’s make a few more figures. I actually originally planned to do like 10 slides and walk you through it. It ended up being 50 slides, but I hope this really explains how this watermarking technique works well, how watermarking itself can fail or be removed, and so forth. So I think it might be an interesting topic because a lot of people use LLMs these days and also consume a lot of text on the Internet that might be generated by LLMs. And now there’s going to be this watermarking, and there’s this, I guess, fear of watermarking making text worse, or what’s actually the benefit of this watermarking? And so what does it mean? And I think if we understand a bit better what watermarking is, that goes a long way, and then we can make up our own minds about whether that’s a good thing or not, and so forth, like the pros and cons. So, my goal here is really to explain how the underlying mechanism works and how they are going to implement this type of watermarking, text watermarking. Slide 2 of 52, time stamp 1:41 It’s also a great example to illustrate why understanding things from scratch is actually quite useful. This watermarking technique is also a nice way to explain how conventional models or LLMs in general work under the hood. So yeah, you may know I like doing things from scratch. Like, I have my books: Build a Large Language Model From Scratch, Build a Reasoning Model From Scratch. I have some articles labeled from scratch. So, for me, “ from scratch often includes coding. So this one will not be coding-related, but coding from scratch is actually a very, very useful technique because it really helps you understand how something is implemented. And then from that we can derive our understanding, figures, concepts, because if we don’t really implement things, if there’s no code, it’s really sometimes ambiguous. And of course, you know, as I realized, not everyone is coding from scratch anymore. Like back in the day, coding something from scratch was all we had. I mean, there were only humans coding. Nowadays, coding can be done by LLMs. However, that doesn’t mean reading code is no longer useful, because it carries a lot of information. So in this case here with this watermarking, spending some time coding an LLM from scratch really makes you realize how this sampling inside is implemented. We still have some relevant code snippets. And then that really, in turn, helps us understand, oh, the watermarking is applied at this position, and this has so-and-so consequences and so forth. So I think even though people may not be coding from scratch, at least not all the time anymore, it is still useful being able to, let’s say, build something from scratch for educational purposes to understand something deeply and then also for research purposes to manipulate this in a transparent way that is not hidden away in tons of layers of abstraction. But that aside, I think it’s just a coincidental nice relationship here because, for this slide deck, I actually used a lot of figures from my from-scratch coding materials. Slide 3 of 52, time stamp 3:58 So a few days ago (this is August 14), there was this article, How Claude’s Text Watermark Works, and there was this article here; it’s just like a screen recording, so it can have everything in the slides, but there’s plenty of detail. They updated it actually a couple of times, so originally when I read this, it was a way shorter. Still, it is very, I guess, conceptual; there’s like this overview, and there’s, I mean, there’s not a single figure in there. And so it’s kind of still hard to understand what they’re trying to do. So they explain a lot about why they’re going to do it, but they don’t explain how. They’re linking to one paper somewhere there, which is very technical also. So I do think it makes sense maybe to take a step back and start at the beginning to kind of understand what they’re trying to implement here with this watermarking technique. And so the motivation, by the way, of watermarking is for them to identify if someone posts some text that they can say, oh, this text was generated by our Claude Opus 4.8 model, for example, so that they have a way to tell, OK, this text is AI-generated because it carries this watermark. And this watermark is invisible to users, so only they can decode it and find out whether the text has their watermark. Why can only they do it? We will get to that later in this (hopefully not too long a video), but one thing at a time. Slide 4 of 52, time stamp 5:35 So I wanted to start with a brief prelude to explain how text generation works in LLMs, because based on that we can then more easily understand how the watermarking works and that this is actually not a huge, expensive thing on top of it. It’s really just like a minor, I guess, tweak inside the regular text generation process. Slide 5 of 52, time stamp 6:01 So when we are using something like ChatGPT, for example, let’s say I ask the question, the capital of Germany is, and yeah, ChatGPT or other LLMs, so this is just like an example would, for example, answer “Berlin”. So here, in this case, it’s generating two tokens, like “Berlin” and the period. But for simplicity, let’s assume it’s generating one token. So the next token is the “Berlin” token. How is this token generated internally? What is happening under the hood when we type something here like the capital of Germany is and receive a token like “Berlin” back? What is actually going on there behind the scenes? Slide 6 of 52, time stamp 6:41 So in the next couple of slides, I want to briefly talk about what happens under the hood when this next token is generated. Slide 7 of 52, time stamp 6:50 So assume again that our prompt is the capital of Germany is. And the first step here is to convert this into token IDs. So tokenizing it and converting it into token IDs is one of the main steps at the beginning. This is outside. It’s not inside the LLM; it’s outside of the LLM. So we are simply converting the text into token IDs. It’s just a format that embedding layers can work with. Slide 8 of 52, time stamp 7:22 And then this passes through the LLM. And the LLM gives us a score distribution for the next token. Slide 9 of 52, time stamp 7:31 So again, this is just like a brief overview of how LLMs work internally. So I’m not covering the LLM machinery itself. I talked about it many times in my other From Scratch LLMs videos and books. The important part is that when we generate the next token (for example, “Berlin”), we have, at this point, a distribution of scores. So this is the output produced by the LLM. Here in this case, we’re looking at logit values. So these are just scores from minus infinity to plus infinity, like a range of scores. Here’s an example, ranging from about -8 or -9 to 20. We could convert these into a probability distribution, but technically, it’s not strictly necessary depending on how we sample. But so you can think of the logit values as the raw scores. And the raw scores go over the entire vocabulary. Slide 10 of 52, time stamp 8:39 That means every possible word that the LLM could generate. Now here, in the vocabulary index, a certain value (index position 19,846) receives the highest score. So I spread out the distribution. If you would run this prompt through an LLM, you would even see something more extreme: that everything is, like, very, very, very close to zero. And “Berlin” would probably be much, much higher even. But just to show you a few, you know, like peaks here so it looks a bit more interesting, I kind of zoomed in; in and spread out the distribution a bit. Now here, “Berlin” is the highest score because you can think of it as the most, I guess, probable or plausible next token if I have a very specific prompt like this. So the other ones, I mean, it could be something like Hamburg or Munich that the LLM might guess incorrectly. But nowadays an LLM should be fairly certain that “Berlin” is the correct answer here. You are also seeing here the vocabulary index. So that’s like over the whole vocabulary. Nowadays, LLMs have like 250,000 possible tokens as output. I’m truncating it here from 19,800 to 19,900 because there’s just so much space here on this slide. If I would have a very realistic vocabulary of 250,000 words, everything would be so narrow that we would barely even be able to tell or see anything on this distribution. So this is just truncated for educational purposes. The important point is that in regular text generation, we get this score distribution. Now, what we do is look at the highest score. Slide 11 of 52, time stamp 10:33 I will get into more detail later on how this is selected. So it’s not necessarily precisely the highest one, but for simplicity, assume we are taking the highest score here. And in this case, it’s 19,846. Slide 12 of 52, time stamp 10:52 And this score is then detokenized, and we get “Berlin” back. So that is the process here on this slide: from an input prompt to conversion into token IDs and tokenization, passing it to the LLM, getting this score distribution, getting the next token, and converting it back into text. Slide 13 of 52, time stamp 11:12 And then this text is appended to the input. So if we have a question that requires multiple output tokens, we keep going in this loop until the answer is complete. That usually means that the LLM generates an end-of-text token, for example, here. For simplicity, I’m showing you only one iteration where it generates one token. But yeah, as I said, it would kind of continue like that, where we are feeding back the modified input to the LLM for the next round. Now, how do we actually sample this next token here? Slide 14 of 52, time stamp 11:44 I briefly said, well, we could just technically select the highest one, the one with the highest score. This is called greedy decoding. That’s one way to do it. But most LLMs, like if you use them, they don’t do greedy decoding where they always pick the highest one. Because if you ask it on some other prompt, it might not be what we want to always have the highest score, because then it would memorize the training data. It would always kind of give the same response and so forth. So we actually often want some variation in the outputs, but not so much that it generates random stuff. So how it works is that, when we sample here from this distribution, we first typically convert it into probability scores. Slide 15 of 52, time stamp 12:28 So here I just have these scores shown in this plot. I’m just using NumPy for simplicity; whatever tool you use (e.g., PyTorch), the same concepts apply. But let’s assume we have the scores here in NumPy. So what I would do is I would compute the softmax. Technically, I would use a softmax function implemented in Torch or PyTorch, for example, that is numerically stable for both large and small values, including very high positive values, very low positive values, and very high negative values. Here I’m just writing it out like that. That’s the canonical softmax, just to make it a bit more readable. But the details don’t matter here. Slide 16 of 52, time stamp 13:20 What matters is that after this conversion, the scores here, I mean, there’s only so much space on the slide, but the scores here, they would add up to one. So it’s essentially like a renormalization. So they would be normalized to sum up to one. That’s all that the probability conversion does: the softmax conversion. So then once we have these probabilities, we can use a random number or, like, a random sampling algorithm. For example, here in NumPy, we could use the choice function or method. So this is with a specific random seed we are passing to the vocabulary indices. And then, and that’s the important part, we are passing the probabilities as the weights. So, these, essentially, yeah, are like: “How likely is a certain token to be selected?” So, for example, if “Berlin”, after this normalization step, the softmax step, has a 99% probability and the other ones together have a 1% probability, then if we would sample 100 times, 99 of the times, we would get “Berlin”. In realistic LLMs, for example, that are well trained, “Berlin” might receive a probability of 99.999999 or something like that. So you’re almost certainly always sampling “Berlin” because it’s very confident that the answer is “Berlin” in this particular case. So yeah, that is how we would sample from this distribution. There are modifications like top-k sampling or top-p sampling where, let’s say, just for simplicity in top-k sampling, we would select the top 100 tokens and then apply this random choice only to the top 100, the 100 highest-scoring ones, so that we don’t get nonsense tokens in there. For this example, it doesn’t really matter. I mean, it’s just like another thing to explain, so I’m skimming over this. So you can maybe assume that this is already the top 100 tokens using top-k or something like that. Slide 17 of 52, time stamp 15:37 And so, for example, here’s an example. If we sample 10,000 times with a probability of “Berlin” being very high, 99.9, we would sample “Berlin” 9,997 times, sample the word “Hal” twice, and one “Moh”. And these are basically nonsense tokens. It rarely happens that, in this case, the LLM might produce nonsense because, as I mentioned before, I spread out this distribution a bit to make it more interesting. A real LLM would probably, 10,000 out of 10,000 times, sample “Berlin” because the probability of “Berlin” is so high. But this is for illustration purposes. Slide 18 of 52, time stamp 16:21 Now we briefly talked about how LLMs work under the hood, which I think is kind of an interesting concept in itself. But I’ve talked about this many times before, so I don’t want to bore you. I just wanted to set up some context for now, explaining how this watermarking works. Slide 19 of 52, time stamp 16:41 So, we mentioned that we select the highest-scoring token when sampling. Or we use this probability sampling, which will lead to one of the highest-scoring tokens being selected most of the time. Now here’s another example without watermarking. I changed the prompt slightly. Now the prompt is: today’s weather is “cold,” and a possible answer could be, for example, “gray” or “overcast”. So in contrast to the “Berlin” example, I would say “gray” and “overcast” kind of are interchangeable. They are both reasonable next tokens for this prompt, given the goal of completing this text or writing the next token. So it’s almost like a coin flip which one we want to select. There is not really an objectively worse one of one or the other. So when we do the random sampling, because they also have relatively high scores and their scores are similarly high since they are both plausible tokens, we might get one or the other. So almost half of the time we would get “overcast”, and almost half of the time we would get “gray” if we repeat the sampling multiple times. And that’s how LLMs often end up with different answers if you provide the same prompt. If you use the same prompt and you ask the LLM multiple times, you often get slightly different answers. And that’s because at certain positions, two possible tokens are almost equally likely, so it will choose one or the other. And that token would then influence all subsequent tokens, and so forth. Slide 20 of 52, time stamp 18:25 Now, I wanted to briefly talk about random number generation. So, for example, if we use a random number generator like this, it will generate a random sequence of numbers. If I run it again, the sequence of numbers is different here. So you can see every time we produce five numbers, they are different. If I set the random seed here, like one, two, three, and I run this multiple times, we still get random numbers, but they are now all the same, right? So they are still random. If we use a random seed, we still get random numbers that are different from each other, but they are reproducible. So whether we use a random seed or not, we still get random numbers. But with a random seed, we get a reproducible sequence of numbers. So keep this in mind: this is just like a little primer, and we will use this concept in a few moments. Slide 21 of 52, time stamp 19:26 So, for example, I mentioned before that we might get either “gray” or “overcast” if we randomly sample. Now, if we use a specific random seed like 42, we would always, for example, select “overcast”. I mean, it’s still a random selection, but we make it deterministic. In this case, given this prompt, the model will always select “overcast”. Slide 22 of 52, time stamp 19:49 If we use a different random seed, the model might select “gray”. Every time we sample, it will always select “gray” as the next token. So it’s still random sampling, but we are making it deterministic based on the random seed. Slide 23 of 52, time stamp 20:04 So, in watermarking, Claude watermarking is kind of like the idea that it sets a random seed. But this random seed, instead of being like a number that is fixed based on, I don’t know, someone writing down a fixed number, they’re using a secret key that is essentially like an API key, a secret key, and from that key, together with the four previous words, they derive this random seed essentially. But the idea is that if I go back one slide, it’s the same as here: there’s essentially a fixed random seed, and that random seed always selects the same next token. Okay, so instead of using random seed 99 here, for example, they have a secret key and also use information about the previous tokens to derive this random seed. But more on that later. Slide 24 of 52, time stamp 21:06 So the idea is that watermarking makes the text generation more deterministic in certain positions. So, for example, if we have these plausible texts on the left side. So if I have a text that says, > The weather today is cold and I may either pick “overcast” or “gray”. And the next sentence could be, > and then “light” or “gentle” They’re both interchangeable again. > And then breeze is “moving” or “blowing” through the trees, and the streets seem “quiet” or “still”. Which means basically I could say either “quiet” or “still”. So there are certain positions in the text where we have token choices where they are almost equally likely, like we have seen before. So that means if we are, this is without watermarking, if we are running the prompt, or given the prompt through the LLM, we might sometimes get this answer here, sometimes this answer, and so forth. And based on the number of positions, we might have 128 possible answers here. And of course, the longer the text, the more positions we have where we can have terms interchangeably, the more combinations, or the more output texts, there are. So, for example, again, one possible output text could be > The weather today is cold and overcast. A light breeze is moving through the trees, and the streets seem quiet. I think I’ll stay home and read a book with a cup of tea. So that is one possible text. Another possible text is > The weather today is cold and gray. A gentle breeze is blowing through the trees, and the streets seem still. I think I’ll stay inside and read a novel with a mug of tea. By the way, it’s also actually raining outside. I don’t know how good this microphone is, but it’s kind of a very fitting context here. But yeah, the bottom line is that you can see there are two very reasonable texts here being generated, and there are more combinations. So they are all reasonable. There isn’t one that is necessarily better than the other. They’re just, you know, slight variations. And if we don’t use watermarking, we might get either one, or it’s just random, right? Because of the random sampling, we might get one or the other. Slide 25 of 52, time stamp 23:31 Now, if we fix the random seed, as I mentioned before, for example, if the random seed is 99, we might always get this text here. So, using a random seed, we can kind of fix which answer we get, because then the random sampling is still random, but it’s deterministic in the sense that it’s reproducible. It’s always going to be the same then. Okay, so that is still without watermarking, now with a random seed. Slide 26 of 52, time stamp 23:59 And the watermarking is essentially doing the same thing. Now, instead of just using a simple random seed, they have a so-called random key, where this random key is involved in selecting the text, essentially. But what we can already say is that, in the Claude blog post, they say the watermarking shouldn’t make the text worse. If we look at this mechanism, yeah, it makes sense why it would not make the text worse. By the way, I’m not defending watermarks here. I’m just trying to explain. So please don’t kill the messenger here. But what I’m trying to say is that the watermarking is nothing else for the end user than fixing a random seed and making this sampling kind of deterministic, if that makes sense. Slide 27 of 52, time stamp 24:47 Okay, so the summary so far is without watermarking. We often sample without a random seed because I know most people don’t even use one. I honestly don’t think you can necessarily do it with the Claude and OpenAI APIs. I know you can do it in Ollama, but I also always had some problems with that because I used Ollama in one of my books for the bonus material to generate some texts. I was fixing the random seed, but it still wasn’t always deterministic, and so forth. So it’s tricky. Your mileage may also vary, depending on the software version and so forth. Anyways, so without watermarking, we have this random sampling. With watermarking on the right-hand side, we still have the random sampling. But in addition to just a random sampling being fully random, we have this watermarking key. And this watermarking key is passed to the random seed generator to set a specific random seed, making this deterministic. But it’s essentially very similar, and like I mentioned, there’s a lot of benefit in terms of understanding things from scratch. And now we know essentially where this watermark is applied to. So this is essentially applied to the sampling. It’s not applied inside the LLM, which is actually cool knowledge. So they don’t need to train a new LLM for that. They can just use an existing LLM, and they just apply it at this sampling stage. They don’t have to retrain anything or anything like that. So yeah, that is actually interesting, right? Slide 28 of 52, time stamp 26:21 But we are not quite done yet. I would also like to talk about how we can understand or see whether text is watermarked. So detecting the watermark is only possible if we have access to the key. So, for example, if we have these different texts, and essentially, after the text was generated, you find some random text on the internet (for example, you find this text number four here on the internet somewhere), you want to know: is this watermarked? Well, it’s impossible to know because, in order to know, you would need the watermarking key. You need this scoring function, and then you have to score basically the text with a scoring function. And then the idea is that if the score is above a certain threshold, then the text is watermarked. Otherwise, it’s not watermarked. But as the end user, we can’t do this because we don’t have this key. So the key is not available to us. Only Anthropic will have the key. However, in this blog post, they mentioned that they are providing it, of course, or they’re going to develop an API for that that they will make available. I don’t know. Honestly, I’m not affiliated. I don’t know the details. I was just reading this in this blog post. That’s all I know. So that API might as well be private for some companies, like, let’s say, X or Substack Notes, when they want to label AI-generated posts. They may make it public for end users to use. Who knows? We will have to wait on that. But yeah, so the bottom line here is that watermark detection is only possible if we have this watermarking key or, of course, the API that they are going to develop. Slide 29 of 52, time stamp 28:03 Now, removing the watermark is interesting. So now that we know how the watermarking works, we also know the shortcomings. I mean, this is really highly dependent on specific tokens in certain positions. So, for example, in this given text, if these colored words or tokens are the watermarking positions, we know that we could remove the watermark by editing this, right? If we change all the words at these positions, we would be 100% able to defeat this watermark. Now, the problem, though, is that we don’t know, right? Slide 30 of 52, time stamp 28:41 So we don’t know where these words are because we haven’t generated the watermark. So we don’t know which positions to look at. So the practical scenario here is that we could just randomly edit the text. So we would randomly change a few words and hope that we change enough positions to edit the watermark. So that would be one way to remove it. And since we also don’t know which are the highest-scoring ones, because that would require us to have access to the LLM and rerun the prompt through the LLM to find out which words are the highest-scoring, we can kind of only guess. So for example, we might say, oh, we replace “overcast” with “cloudy” because we don’t know that “gray” was high-scoring, you know? So in this case, it might be intuitive to say “gray”, but there might be cases where it’s not so intuitive. So what I’m trying to illustrate here is just some general text editing where we are modifying positions, but we are still kind of guessing what a watermark position is. So since we don’t know, we added just a few words here and there. And if we added enough words, that would also defeat the watermark. Slide 31 of 52, time stamp 29:59 So yeah, that was the watermarking in a nutshell. I mentioned that there is a scoring function to find out whether something is watermarked. And I want to do it as a bonus here. It’s already a long video, but as a bonus here, I wanted to briefly also explain how this scoring function works because that is also interesting information. It’s a bit complicated. It’s not essential to understand how the scoring function works. But the reason why they do it the certain way they do is to make the detection cheaper. Because otherwise, if I go back one slide or two slides, if you wanted to check if something is watermarked, if even they wanted to check, they would have to rerun the prompt to get these scores and then apply this watermarking random seed to get this text and then compare. And that would be very expensive because then essentially every text you want to compare, you would have to rerun the LLM. You have to know which LLM, and that would be really unfeasible because you often also don’t even know the prompt, right? So yeah, so they have like a trick that they use to, yeah, I would say, modify the sampling so that you don’t use or don’t need the LLM later on for the scoring stage. And in the blog post, they mentioned that they derived this method from a paper. It was a Nature paper, and this method is called SynthID-Text. So that was like a paper that came out maybe one or two years ago. It was by Google, and they use a similar technique they call Claude watermarking. I don’t know, sorry, I don’t know if they use exactly that technique, but that’s the one they mentioned. Slide 32 of 52, time stamp 31:39 So how does it work? So before we looked at the slides, we looked at the regular, let’s say, overview here, where we have some text. We put it through the LLM. We get this logit distribution and then we sample from the distribution and get the output token. And here, during the sampling, we use the watermarking key and the random seed generator. So this is still correct. This is still what’s going on, but there is a bit more, I guess, nuance to how this token is sampled. So they’re not just using, let’s say, NumPy’s random choice. They’re using something a bit more sophisticated here. Slide 33 of 52, time stamp 32:16 So assume, again, our context is “the weather today is cold,” and we want to generate the next token. So, for example: “gray”, “overcast”, “gloomy”, “cloudy”. “Gray” is 50% probability, “overcast” is 30, “gloomy” is 15. Let’s say “cloudy” is 0.05 and the rest is, let’s say, 0. Here it looks, of course, a bit different. Let’s say that’s “gray” and “overcast”. I’m just reusing this figure. But now imagine these are the most likely ones, like “gray” and “overcast”, and everything else is just very small, except “gloomy” and “cloudy,” maybe. So essentially, think about just a very small vocabulary for this example of four words instead of all these 50 words here, just to make it even simpler. Now, as I mentioned before, we could use ‘sNumPy’s random choice with these probabilities to sample the next token. Slide 34 of 52, time stamp 33:13 And we could use the watermarking key with this random seed generator to make it deterministic and get the certain watermark that we want. But as I mentioned before, this would be very expensive. Not the sampling itself. That doesn’t matter. This is pretty cheap. But the detection later on would be very expensive if we are trying to check random text on the internet. Slide 35 of 52, time stamp 33:35 So instead, what they use, they also use it during the sampling, during the generation, so that it can be reused later during detection. What they use is called tournament sampling. So this is instead of using something like random choice, they use a concept called tournament sampling. And so how does that work? It might look a bit complicated, but it looks really more complicated than it really is, to be honest. So you might have to, I guess, stop the video at some point and just sit with the figure a bit. But I think it is actually simpler than it looks like. It’s like once you get the hang of it, it’s pretty straightforward. But let me try to explain here. So what we have is we have still this context, and then we have these probable or plausible next tokens with these different probabilities. Now they have something they call random watermarking functions. Slide 36 of 52, time stamp 34:35 Here we have three watermarking functions, G1, G2, and G3. In reality, they might have 30, 50, or even more. Here I’m just using three because that is simpler on this slide. It’s just smaller, you know, like it fits better on the slide. Now, if we look at this word “gray”, this might give us a signature 101. With that, I mean, if we use this watermarking key to generate this random seed, and we have three functions, G1, G2, G3. If I put the word “gray”, what I’m skipping here is that usually you put the word “gray” together with the four or three previous words from the context. So it’s “cold” and “gray”. If I put that into G1 together with this watermarking key, I get the value one. Why? Well, that’s just how this function works. It’s like a random function. The random function either returns zero or one. In this case, with this random key and this token, it returns one. With the same key, but a different function, you get the value zero. And then here you get a one again. So if we have more functions (of course, 30 functions), this will be a very long string of ones and zeros. Slide 37 of 52, time stamp 36:00 It’s basically like a bit string, like if you have bits of zeros and ones. Okay. So this is for “gray”. So we get the signature 101 through using these watermarking functions. Now we can do the same thing for all the other ones. So we can do it for “gray”. We can do it for “overcast”, “gloomy”, and “cloudy”. So each one has a different signature here. So, for example, “overcast” is zero, one, zero. “Gloomy” has zero, zero, one. “Cloudy” has one, zero, zero. Okay. So we have these bits here now. The next step is a so-called tournament sampling where we just pair them. Slide 38 of 52, time stamp 36:39 Like, you know, like a soccer tournament, the knockout (KO) stages, or like the playoffs in American football, you always have two teams playing against each other. And that’s kind of like the same idea. We have a pair of tokens, and they’re playing against each other, essentially. And the scores, they come from these functions here. So we start with the first function in the first round. So we have “cloudy” and “gray”. So we look up here: “gray” is a one and “cloudy” is a one. Okay. So one and one. “Overcast” and “gray”. So “overcast” is zero, “gray” is one. So we have zero, one. “Gloomy” and “overcast”. So here we have “gloomy” zero, “overcast” zero. So zero, zero. And then we have “gray” and “gray” again, because we are running out. So we don’t have enough of the others. So we have one duplicate. So this is chosen randomly. And so you have one and one here. Now we look at the results. So this is a tie. In the case of a tie, we also select randomly using, you know, the random seed and the watermarking key. So here, “cloudy” survives. And from this one, G1 is, according to G1, “gray” is the winner because it has the one. So “gray” survives. And then here, “overcast” and “ gray “ are a tie, randomly selected, and “gray” also randomly selected. So we have now “cloudy” and “gray” and “overcast” and “gray”. And we play the next round in this tournament. So in this next round, we use G2. So according to G2, “cloudy” has a zero here. “Gray” also has zero. “Overcast” has one. And “gray” also has zero, sorry. And so, the next stage of the tournament again. Slide 39 of 52, time stamp 38:24 So we have a tie. We randomly select “gray”. And here we have “overcast” as the winner. And so we have “gray” versus “overcast” in the final. And then we look again at the scores. So “gray” has a one. “Overcast” is a zero. So “gray” is the winner. And that’s how the token “gray” is sampled. What is the watermarking key doing here? So the watermarking key, if I go back a few slides, is selected for generating these scores using these random watermarking functions. So the watermarking key determines essentially what values we get at these stages. So the watermarking key is still very important. Otherwise, these signatures would look different. Slide 40 of 52, time stamp 39:11 So we now have sampled the next token. And that’s just how this modified sampling procedure works. We could have used NumPy’s `random.choice`. But the shortcoming of that is that if we want to score random text on the internet, we would have to rerun the LLM. With this technique, we don’t. I will show you in a moment. So this technique sounds like really weird and cumbersome, but it has the advantage that we can now score random text more easily without having to rerun the LLM. So it’s essentially just to make the detection easier and cheaper. Slide 41 of 52, time stamp 39:43 Slide 42 of 52, time stamp 39:48 So, for example, if we have a new text. So I’m just using the same text here, but let’s assume it’s new text. So this is after the sampling, when we are scoring. And let’s say we are discovering this text on the internet. And the text is the weather today is cold and “gray”, and we want to know if this is LLM-generated or not. So we would, or Claude/Anthropic would, have the watermarking key and these functions: G1, G2, and G3. And it would put this text through these functions. For the one position here for “gray”, we would get 101, similar to what we got during the generation process. So this is the same as before. And this has, if we add up these bits, two bits, right? One and one here. So it has two bits of information, let’s say, for simplicity. This is just a really simple illustration. But let’s assume we get a score of two here for the “gray” in this position. If we had a different word here, “overcast,” in this position, we would get one if we get “gloomy,” like we also have one, and “cloudy” one. So I’m just summing over each row here, right? So that’s just like a score we would get at each position. And here I’m only looking at the last position. If I would do this at other positions, I would get a different score at different positions. So, for example, let’s assume at the first position I get a two. Here I get a two. For “today”, I get a three. For “is”, I get a two. “Cold”, two. And “gray”, three. So here I’m applying these watermarking functions as I’ve shown on the previous slide. Slide 43 of 52, time stamp 41:33 And I’m just adding up these numbers across the three functions. And the watermarking functions are very cheap. So you can just quickly run them on the whole text and get these scores. And then based on that, I can compute the average bits. So if I just average over all these values here, let’s say I get 2.23. Slide 44 of 52, time stamp 41:55 Now, if I have slightly different text, so here I swapped “today” with “now” and “gray” with “overcast”. These now get a score of one and one. And if I average over this whole string, then I get a 1.71. And so for that, I don’t need an LLM. All I need is the watermarking key, the random seed generator, and these functions, G1, G2, and G3. And that’s all I need. I don’t need the LLM. And I can get this score here. And what they do is apply a threshold. Slide 45 of 52, time stamp 42:27 So, for example, I mean, they don’t use this exact threshold. But for example, we can say if the score is greater than two, then the text is watermarked. If the score is smaller than two, it’s not watermarked. So here, if the score is greater than two, it’s a yes. So yes, this is watermarked. In this case, 1.71 is not greater than two. So this text is not watermarked. Okay. So that’s just the way we can then detect whether random text on the internet is watermarked or not. It’s essentially just applying these watermarking functions and then averaging over the scores and applying a threshold. Okay. Slide 46 of 52, time stamp 43:13 So again, the tournament sampling is mainly to make detection easier and cheaper. We could also use something like NumPy’s random choice with a random seed or to make the sampling deterministic. But then again, it would be hard to score any text on the internet. Slide 47 of 52, time stamp 43:30 So yeah, the summary is still the same, though. The thing that is different between no watermarking and watermarking is that we are controlling this sampling here with the watermarking key. And inside that, we have this tournament sampling. And yeah, as I mentioned before, detecting the watermarks requires the secret key and the watermarking functions G1 to Gn. Slide 48 of 52, time stamp 43:54 And again, removing the watermark, because I think that’s maybe interesting to some people, would ideally involve editing all the positions here. But since we don’t know which positions are watermarked and internally, they choose the positions so that they have equally likely tokens at those positions. And there might be positions where that’s not true. So here, for example, for “trees”, we might not even have an alternative word that is high scoring so they don’t watermark that position. So they only do the watermarking at certain positions essentially. Since we don’t know which positions to kind of defeat or remove the watermark, we would... Slide 49 of 52, time stamp 44:29 ...have to edit several places in the text. So what I think that means for the future of AI-generated text is that this actually... Slide 50 of 52, time stamp 44:36 …might result in worse AI-generated text. So I think if there’s a person who likes to use AI-generated text everywhere on the internet, let’s say there’s a news website that likes to use AI-generated text to write the news, I don’t think watermarking will necessarily stop them from doing that. They will probably still want to generate AI-generated text because that’s part of their workflow. So I think my guess is that they’ll use another model. Slide 51 of 52, time stamp 45:08 They’ll just use a second model to edit the text to get the so-called edited AI-generated text. So it’s complicating the pipeline. Instead of getting the text directly from Claude, it’s now using Claude to generate AI-generated text, passing it through a local model, and then having edited AI-generated text, which is likely not watermarked anymore. So why a local model? I just think a local model because I think all the providers- the proprietary LLMs, not only Claude, but also Google— I mean, Google wrote this paper, right? So I’m thinking that they are also watermarking Gemini text. And I think OpenAI is probably already doing it or will do so as well. I mean, I’m just speculating, but I’m imagining everyone will probably do something like that because there’s like an EU regulation that requires that. And that’s, according to the blog post, apparently why Claude is doing it. Yeah, so I’m thinking local models may not, at least not yet, implement this watermarking. So I think people will just use a local model and then generate edited AI-generated text. And my guess is it will be slightly worse than the original text because for the local model, you might now be using a smaller model. So, I mean, you could also technically just use the local model directly to generate text. But in my view, editing text is simpler than generating text. So for the generation of the text, you might use a very expensive high-end, I don’t know, like the highest, most expensive Claude model for complicated text. And then you use a cheaper local model to make these surgical edits, essentially. That’s probably what’s going to happen. And why worse? So if we look back at this graphic where we just added random positions, you might be just changing words for the sake of changing them. And then it risks making the text worse. So you might still have generated text, but it’s kind of like it’s edited awkwardly. Slide 52 of 52, time stamp 47:20 But anyway, so my goal here was to explain how the watermarking works and not, let’s say, the worldwide ramifications of that. But I hope this kind of behind-the-scenes, under-the-hood look is useful. The watermarking is not as complicated as it might seem, but I think it was still 52 slides, so it was also not super trivial. So I hope you found this little lecture useful. And yeah, until next time, see you then. PS: If you like more explainers in this style, I don’t post videos to YouTube regularly, but I have accumulated over 300 videos over the years, which you can find on my YouTube channel here . I also have a YouTube version if you prefer using the YouTube player And here is a link to the slides

0 views
Evan Hahn 2 days ago

Vim's UserGettingBored autocmd

In short: Vim has a joke autocmd called that doesn’t do anything. Vim’s automatic commands feature, usually shortened to “ autocmd ”, lets you run code when various events occur. For example, you could implement an auto-save feature by binding the event to the command. Vim has over 100 events, from “buffer was created” to “file was saved”. But one of them sticks out to me: . Here’s the documentation: : When the user presses the same key 42 times. Just kidding! :-) When I saw this, I was busy doing something else and it completely derailed me. “I must know more,” I thought. Here’s what I found: Unfortunately, it doesn’t do anything. It only exists in the documentation (and some tests). If you try to use it with somethig like , you’ll get a “no such group or event” error. It’s present in Vim , Neovim , and Vim Classic . It was first added by Bram Moolenaar in July 2000 , over a year before Vim 6.0 was released. The original description was, “When the user hits CTRL-C. Just kidding!” And it didn’t do anything back then, so I don’t think it’s ever been real. In August 2001, he added the smiley face to the documentation . It then read, “When the user hits CTRL-C. Just kidding! :-)” Twelve years later, in 2013, the description changed to its current iteration: “When the user presses the same key 42 times. Just kidding! :-)” In 2022, developer Mike Smith created an unofficial plugin inspired by this joke autocmd . If you press the same key 42 times in Insert mode, a picture of Samuel L. Jackson appears. 22 years later, it’s finally real. Unfortunately, it doesn’t do anything. It only exists in the documentation (and some tests). If you try to use it with somethig like , you’ll get a “no such group or event” error. It’s present in Vim , Neovim , and Vim Classic . It was first added by Bram Moolenaar in July 2000 , over a year before Vim 6.0 was released. The original description was, “When the user hits CTRL-C. Just kidding!” And it didn’t do anything back then, so I don’t think it’s ever been real. In August 2001, he added the smiley face to the documentation . It then read, “When the user hits CTRL-C. Just kidding! :-)” Twelve years later, in 2013, the description changed to its current iteration: “When the user presses the same key 42 times. Just kidding! :-)”

0 views
Sean Goedecke 2 days ago

You should never be angry at work

I try not to give a lot of prescriptive advice about working in tech companies 1 . There are many ways to be successful, and every company works differently. If you’re shipping projects and your management chain is happy, it doesn’t really matter how you’ve accomplished it. However, there’s one thing that I do think is solid advice: you should never be angry at work . Anger in the workplace is toxic. An angry colleague immediately becomes a new problem to be managed, not a professional helping you manage problems. When someone is visibly angry in a meeting or in Slack, it kills the entire atmosphere: other engineers will often go quiet entirely, not wanting to make the situation worse. If you routinely “get heated” at work, the best-case scenario is that you’re part of a tight-knit team of confident people who aren’t put off by it 2 . No harm, no foul. But the second someone comes onto your team who’s not so confident, or you have to communicate outside of your team, it becomes a big problem. Healthy workplaces route around anger in the same way that networks route around damage. Emotionally unreliable engineers will get left out of conversations that might cause them to blow up. Decision-making will get done around them in backchannels. I’ve seen this become a self-reinforcing cycle: angry engineers aren’t consulted on key decisions, which makes them angrier, which pushes them even further away from the spaces where decisions get made, and so on. You can often find these engineers bitterly complaining that they keep the company together, but nobody ever listens to them. In my experience 3 , this is almost never true. Engineers who are highly effective tend to get listened to — at minimum by their colleagues, and eventually by managers and product managers who want to extract as much value as possible from them. (One reason this is true is that all successful projects involve working with other people, and if nobody listens to you, you can’t do that.) Why do angry engineers believe they’re important? Paradoxically, anger can be really useful to a software engineer . Angry engineers are rarely the ones holding the company together, but they’re also rarely useless . One surprising thing about working for big tech companies is that some engineers are not just unproductive, but actively net-negative : either because they’re incapable of doing useful work on their own, or because they’re sloppy enough that they create more work than they do, or because they’re so checked out that they literally do nothing. Angry engineers might be net-negative in a cultural sense, but in terms of literally solving tickets and shipping features, they’re usually well above average. Why is this? Anger often comes from caring about your work, and caring a lot is sufficient to make you a competent engineer . I’ve never worked with someone who genuinely cared about their work who wasn’t (or didn’t eventually become) competent. I actually think it’s healthy for an early-career engineer to sometimes get angry about their work, because it means they care a lot: it’s still a mistake in the moment, but it’s a “good mistake” . I certainly used to get angry — in fact, I wrote about the angriest I’ve ever been at work here 4 . But you have to move past it . Think of “caring about your work” as a vertical tube, unsealed at either end. You fill the tube by pumping in emotional investment from the bottom 5 . If you have too little, it drains away and you end up as a useless coaster. But if you have too much, it overflows and you end up as an angry engineer that people have to work around. One solution is to try and care the exact right amount: be invested in work a bit, but also have hobbies and a family and whatever else gives you perspective about your work problems. If you have a rich and healthy personal life, it’s hard to find yourself yelling at somebody about React state management. However, this is a tricky balance to maintain over time. Another solution is to care about different things. The reason too much caring overflows into anger is because what you care about is misaligned with what the organization cares about . If your interests are perfectly aligned with your company’s (for instance, if you primarily care about delivering shareholder value ), you can fit way more emotional investment into the tube before it overflows. Here’s some dangerous advice : showing a little bit of anger at work can sometimes be useful. It can be a good way to signal that you care, or to build rapport with certain people, or to draw attention to something you think is important. However, it’s still always a mistake to be angry. You need to be able to drop back to a friendly mode at will, which is very difficult when you’re genuinely angry. Being able to show a full range of emotion at work is good. It makes you more persuasive and more human. Being a fully professional robot is fine — you can have a successful career this way — but there’s always going to be some kind of uncanny-valley HR-ness to your work persona that will make it hard to connect with your colleagues. If in doubt, don’t show anger. It’s never wrong to be professional. However, if you can signal that you’ve got enough distance to separate your professional feelings from your real feelings, and enough perspective to realize that the stakes of a technical decision are fundamentally not that high in the grand scheme of things, it can sometimes be okay to show visible frustration so that people know you’re still human. Well-known software engineering personalities are often angry. It feels unfair to give too many negative examples, but obviously Linus Torvalds’ rants about Linux are a great example. Some of my favourite engineering talks are from Bryan Cantrill, who is sometimes visibly furious at his subject matter. There are too many well-known angry blog posts to list, but I’ll cite one I genuinely like: my Australian blogging colleague Nikhil’s post titled I Will Fucking Piledrive You If You Mention AI Again . Anger is a part of the general image of a competent software engineer. Many junior engineers learn from this that it’s okay to be angry. However, taking your emotional cues from engineering celebrities is a big mistake, for a few reasons. First, you are not Linus Torvalds or Bryan Cantrill . Torvalds is the BDFL of the most important software system in the world. Cantrill is the cofounder and CTO of his company. When these people are angry at work, people will not work around them, because they are the ones deciding what gets worked on . Once you’re the one in charge, you can get away with being emotional in the workplace 6 . Second, you don’t know what it’s like to work with these engineers . People give talks and write blog posts because they’re emotionally worked up about something. If your only exposure to a celebrity is via their conference talks and blog posts, you’re seeing them at something like their maximum emotional intensity. If you then take that level of emotion into your normal everyday work, you’re almost certainly overshooting. I’ve been reorged into dysfunctional teams, have had projects I enjoyed cancelled, and have worked on systems that were extremely chaotic. I can’t remember the last time I was actually angry at work. To be clear, I’m not successfully hiding my anger (unless it’s so repressed it’s invisible to me as well) 7 . Nor am I naturally a chill person. I’ve just reached a point in my career where I genuinely don’t get upset about work stuff. A cynical person might say here that I’ve stopped caring about my work, so of course I don’t get angry anymore. I’ve left the side of the “real engineers” — the Linus Torvalds and Bryan Cantrills of the world — and sold out for that sweet, sweet big tech money. I mean, maybe! It’s true that I’m less invested in specific technical decisions than I used to be. But I still care a lot about doing a good job, I still spend a lot of time tweaking and reading code, and I certainly get more done than I did when I was more emotionally volatile. Being angry at work feels good. It feels like proof that you’re working on something that matters, and that you’re personally having an impact. If you’re angry, nobody can call you a coaster. But anger is only a local maximum. If you can find your way to a different style of working, you’ll not only be more effective, but you’ll be in a far better place to have impact on problems that actually matter. Mostly, I fail . As I understand it, this is the work environment that most famously angry engineers came up in. My experience is certainly limited (a handful of companies, and maybe ten different teams or organizations). I can certainly believe it happens! About halfway down, in the section titled “it’s not your manager’s fault”. Emotional investment is a liquid with the viscosity of water. To a point. Even Torvalds famously said he’d gone too far with the anger and decided to turn it down a bit. I suppose I’m not the best person to judge whether this is true. If you work with me and I do come across as an angry guy, please do tell me. Mostly, I fail . ↩ As I understand it, this is the work environment that most famously angry engineers came up in. ↩ My experience is certainly limited (a handful of companies, and maybe ten different teams or organizations). I can certainly believe it happens! ↩ About halfway down, in the section titled “it’s not your manager’s fault”. ↩ Emotional investment is a liquid with the viscosity of water. ↩ To a point. Even Torvalds famously said he’d gone too far with the anger and decided to turn it down a bit. ↩ I suppose I’m not the best person to judge whether this is true. If you work with me and I do come across as an angry guy, please do tell me. ↩

0 views

Code-Reviewing My First Monad Implementation

I’ve been wanting to get back into the habit of blogging, and thought a fun project would be to go back through my old code and review it as the programmer I am now. So let’s do that. Eleven years ago, I was so proud to written my first monad that I typed up a big ol blog post about it. This seems like a fun thing to revisit, so let’s do it. To save you the trouble of reading that ancient-ass blog post, here’s the gist. Given a big piece of state, that has many smaller stateful subcomponents, for example, a is full of s: my monad acted like except that you can restrict the piece of the state you’re allowed to manipulate without losing the ability to look at the rest of the state. In the blog post, this monad is called , but at some point in the source code it got renamed to . Cool name. Furthermore, the implementation changed rather dramatically from the blog post, so let’s chase the implementation as given. I’ll reproduce it here: Already there’s a lot to look at. The most immediate thing that strikes me is the formatting; now I’m very much a “use two spaces for indent” kind of guy — Haskell doesn’t provide many opportunities for natural linebreaks, and our horizontal space is much more limited than our vertical space. So don’t throw away your horizontal budget on initial spacing. It’s minor, but being an artisan means caring about the minor details. Much nicer. The other exciting thing to see in the earlier snippet is that this code comes from before the Functor-Applicative-Monad transition. Which is to say that it’s from the long long ago before was a superclass on . There’s some history for you. And then there is the elephant in the room. Whatever this is being passed around by , the instance branches on it and returns when it’s . What the hell is that???? Digging into the instance provides some insight: So returns and some value, presumably which is there to try to warn someone that there’s an floating around. This is a stupid design. If you’re ever in the business of needing to store data-which-might-not-be-valid and a separate tag stating whether or not it’s valid, a much better design here would be to just use instead of . In fact, perhaps I realized this later on, because does exactly this logic: Let’s roll back a bit. The pivotal primitive of is the aptly-named : Pretty reasonable definition here. I’m not sure what the idea behind being a flipped version of was for. The only odd choice here is that in the big tuple returned in , gets passed along unchanged. We can tell from looking at it that here is the current lens we’re looking at. But is the only primitive that changes the lens, and it takes a as an argument. Which is all to say that there isn’t anything actually stateful going on with the . So without looking any further, I’d bet dollars to donuts that this getting returned is purely vestigial and serves absolutely no purpose. Which in turn means we should be able to chop it out of the definition of . Speaking of stupid anti-patterns, here’s another one: This is horrendous for a few reasons. is an unsafe lens that will crash if you try to read out of a . But then that behavior is guarded in by checking that it isn’t already nothing. Stupid. The last function I want to point out is : which… does… something. I guess it attempts to run a -valued , and if it fails, unwinds the state and returns some default . Rearing its ugly head here, however, is the ghost of the nonsense from . Note that if we invoke we will get a crash rather than the roll-back-and-default behavior that it promises on the tin. So, all in all, I am not particularly impressed with my first monad. Here’s how I’d write it today. First, I don’t think I’ve written a “control”-y sort of monad from scratch in the last five years. Whatever behavior you want is almost always just the composition of a few monad transformers. In this case, we have a for the current lens, a for the original state, and a to give us semantics. Since we want our state to roll-back when we take an alternative path, we must put underneath our . 1 This gives us the nice formulation of a as: (we use here instead of since the latter requires . Which happens to be turned on in the original implementation!) Once nice thing about defining our monad as a over a series of monad transformers, is that it allows us to newtype-derive all of the relevant instances: No implementation for such things is required, and therefore we know that the derived instances must be correct. Which means we have many fewer things to worry about. I prefer to use as my naming scheme for these newtypes, so that I can provide as a more user-friendly function. What’s also nice about using standard monad transformers for your implementation is that you can reuse all of their existing combinators. For example, to implement all we need to do is invoke , which lets us change the type of the reader. Notice here that I like to minimize my points without going point-free; as written is very legible, but writing it entirely pointfree is a crime against humanity: Rather than implementing directly, we can note that it is the composition of two pieces — (1) attempt something, (2) return a default if it failed. Whenever you notice de-composition opportunities like this, take them! We can write , which is wildly general and doesn’t care one fig about jails or the legal system: and then is merely All in all, significantly nicer if you ask me. But I’m just a TV. Don’t take my word for it. Compare for yourself! Maybe this was interesting to you. It certainly was fun for me, because I like tearing apart bad code, and I don’t get to see much of it now that I have very talented colleagues. Thankfully I have a huge amount of crap code from my past, including an even more egregious monad that we’ll tackle next time around. Unwrap the definitions of and to see why. ↩︎ Unwrap the definitions of and to see why. ↩︎

0 views
Brain Baking 3 days ago

Deconstructing A 17th Century Letter

Continuing the letter writing fever, I’ve been exploring the academic work called Letters As Loot by Gijsbert Rutten and Marijke J. van der Wal. It’s a treasure trove of information on how letters were written in the 17th and 18th century, based on a corpus of 50k letters that were captured by raiding ships between England and The Netherlands. These boats usually contained a bag full of letters headed to offshore family and friends but due to the raging Anglo-Dutch War (four of them between 1652 and 1784), many letters never made it to their destination. That’s frustrating for the reckless explorers waiting to hear something from their family back home (or the other way around). But at the same time, nearly four hundred years later, these lost letters are the only thing that help us shape our view of lower and middle-class Dutch letter writers, which is otherwise dominated by upper class male writing. I’d love to think that if the Belgian Post loses a letter I sent, it’ll turn up in a couple hundred years for researchers to marvel at. I’m pretty sure by then this website is long gone. Who says the internet will still be alive? Or the planet, for that matter, judging by the extreme weather conditions… Unfortunately, Letters As Loot is written in a stereotypical form of Academese and hard to recommend in its current state. So I scanned the pages diagonally and summarised the coolest findings here. The most interesting is the deconstruction of a typical 17th century letter, written by Angenieten Cornelis to her husband at sea Roelandt IJosten Ostvorndijck. Here’s what she originally wrote: A letter by Angenieten Cornelis, dated 15 September 1664. Lifted from Letters As Loot, page 80. That’s quite challenging to read even for a Dutch speaking person like me. Here’s the line-by-line translation by the authors: What’s so interesting about this letter? The fact that more than 50% of its contents is what we’d nowadays call formulaic nonsense: If you’re married to Roelandt (written as Roellant, which I’m quite sure is misspelled), why write such a cold an distant contents? She’d “regret to hear it” if he wasn’t in good health anymore. I think my wife would kill me if I’d write like that, thinking there must have been a very suspicious reason for my aloofness. She must have made up for that by ending it lovely: a hundred thousand times good night by me —but just to be sure, write your full name again. It is I, Leclerc! Some expressions, such as a way of ending the letter, were so common that they were abbreviated: uwe dienstwillige vrient (your faithful/obedient friend) would become U.E.D.W. . The most funny one most certainly is ick sal onse kijndt eens soenen voor u (I’ll kiss our child for you). It only appeared a couple of times so the writer must have been a creative one. The use of a very formulaic structure was extremely common in the 17th century, which gradually disappeared in the 18th century and onwards. The authors of the research also discovered that this greatly varies between classes: lower class and female writers are much more likely to write like this, while upper class and male writers seem to skip this entirely. It hadn’t occurred to me that this is largely due to (il)literacy. Letter writers having difficulty writing will stick to copying example letters from letter-writing guides. The authors make the distinction between more elite manuals, the very popular school books, and Jacobi’s Gemeyne zeyndt-brieven from 1645 that was reprinted well into the 19th century. Most of these books came with example letters and offered formulas for starting a letter, expressing regret, how to end it, how to note the date, and so forth. Kloek en gezond (strong and healthy), laat u weten dat (to let you know that), blijft hiermede den here bevolen (stay with this the Lord recommended) and so forth are all examples from Jacobi’s book 2 . If we look back at Angenieten Cornelis’ letter, you can easily discern some basic structure and delete all the junk to filter out what Angenieten actually wanted to tell her very respected and devoted husband. Here’s how I imagine that letter might be written nowadays: How much formulaic sentence and paragraph structures do we use nowadays? Barely any, and yet, the main form of the letter hasn’t changed since the 17th century: there’s still an opening (date, addressing to), a short formulaic exchange of pleasantries ( hope you are well, how’s things ), the bulk contents of the reason for sending the letter (here’s some stuff), and a closure ( cheerio or as M.I.A. would write, XXXO ). As the time went on, the strict use of formulaic structure disappeared. Thank God for not having to thank God a thousand times good night anymore. Most letter writers, including me, would start a letter with the date & location, but most older letters would switch this up and end with it. Also, the address we now place on the envelope would be the beginning of the letter, integrated as part of the text. This, together with the formulaic weirdness, would make it easy to spot if a piece of paper with scribblings on it was in fact, a letter, and not some customs clearance notice, tax slip, or worse, declaration of Independence. In sum, for the Dutch history nut into letter writing, this bottom-up research of the “ Sailing Letters ” by Rutten and van der Wal is a gold mine. I didn’t even touch the importance of these letters to the etymological history of written versus spoken words in Dutch. The digital version is freely available. Don’t let the dry text discourage you; just scroll through it to find a cool photo and read a few bits here and there. Time to buy a good brown ink. Looft God Booven Al! Did you know that the address field didn’t need a number? “To cobbler Jeff, the one living near the town square” would suffice. Or “To my husband Jeff, at ship The Jolly Jumper at sea left via Amsterdam commandeered by captain DoGood”.  ↩︎ Later authors attempted to modernise Jacobi’s increasingly outdated manual, but as Rutten and van der Wal note, the “differences outweigh the similarities”: most typical formulaic sentences were copied from other letters, not from examples out of manuals.  ↩︎ Related topics: / letter writing / By Wouter Groeneveld on 21 August 2026.  Reply via email . God is praised in the “address field” 1 ? Why write worthy and very beloved husband to address and start the letter? Why repeat the full name of your husband and yourself? Surely your husband will know your name? Nothing more on this occasion than commend to the Lord? Did you know that the address field didn’t need a number? “To cobbler Jeff, the one living near the town square” would suffice. Or “To my husband Jeff, at ship The Jolly Jumper at sea left via Amsterdam commandeered by captain DoGood”.  ↩︎ Later authors attempted to modernise Jacobi’s increasingly outdated manual, but as Rutten and van der Wal note, the “differences outweigh the similarities”: most typical formulaic sentences were copied from other letters, not from examples out of manuals.  ↩︎

0 views
Stratechery 3 days ago

2026.34: App Snore

Welcome back to This Week in Stratechery! As a reminder, each week, every Friday, we’re sending out this overview of content in the Stratechery bundle; highlighted links are free for everyone . Additionally, you have complete control over what we send to you. If you don’t want to receive This Week in Stratechery emails (there is no podcast), please uncheck the box in your delivery settings . On that note, here were a few of our favorites this week. This week’s Sharp Tech video is on the turnover at DeepMind. Apple Makes Compromises in the EU. Ben has covered the angst surrounding the App Store since the beginning of Stratechery and was focused on Apple’s policies long before it was cool. Now that the company’s finally been forced to compromise in various forums — including a settlement this week with the EU, as well adjustments to its ATT policies in Germany — I thought the most remarkable aspect of Ben’s coverage on Wednesday was how incidental and boring it all seems in the shadow of the possibilities and concerns that exist everywhere else in tech right now. We had a fun conversation about that dynamic at the top of this week’s episode of Sharp Tech before turning to AI cybersecurity, vibe coding epiphanies, and more insight on writing with and without AI. — Andrew Sharp Truth (Social) and Reconciliation. Sharp China returned from its annual August hiatus this week, and in an episode that’s outside the paywall , we talked about various sources of U.S.-China friction before Xi’s visit to D.C. in September. Before that, however, we began in Korea with more questions than answers as Foreign Minister Wang Yi descended on Seoul in the wake of President Trump’s abrupt Sunday evening decision to reduce joint military exercises between the US and ROK. As for that Trump decision, in this week’s Sharp Text article , I used the Korea news as an opportunity to marvel at the exhausting economy of takes and theories that accompanies every foreign policy decision (and meme) under the current administration. — AS August Fun with the Clippers and Lakers . During the quietest period of the NBA calendar, there’s actually been quite a bit of news out of L.A. On one hand, we have a terrific mess as Buss family members squabble and Mark Walter’s DOJ-flavored cashflow problems have led to a shocking sale nine months after he initially purchased the team. On the other, Steve Ballmer and the crosstown Clippers might be in the (relative) clear after a 12-month NBA investigation into alleged salary cap circumvention. We discussed all of it on this week’s Greatest of All Talk , including frustrations with Clippers media coverage, what the NBA wants for the Lakers, and a memorable Top 5 segment about our top vacations.  — AS Stripe Acquiring OpenRouter, Aggregating AI?, Flipping the Business Model — Stripe is reportedly acquiring OpenRouter, an implicit bet on a future market of models and the chance at Aggregation. Nvidia Backs OpenAI Data Center, Anthropic News, Google Buys Spirit Airlines Data — Nvidia makes another deal, this time with a frontier lab; Anthropic’s revenue continues to amaze; and maybe data finally is oil. Apple Settles With E.U., U.S. App Store Fees, ATT Rules in Germany — Apple’s App Store is finally facing the reality of lower fees, and the EU should be satisfied with its work; it’s ok it’s late. So What Was Trump Saying to South Korea on Sunday? — A snapshot of Truth Social foreign policy and the take economy it inspires. More on Watermarking Apple Settles With EU How TSMC Uses Old Fabs to Make New Chips China Built 700 Waste-to-Energy Plants in 6 Years Wang Yi Visits South Korea; Remembering Zhu Rongji; US-China Ahead of Xi’s Visit; How China Monitors Foreigners August Fun with the Lakers and Clippers, Top 5 Takeable Teams or Players, Top 5 Vacations The App Store in the Shadow of AI, Offensive and Defensive Cybersecurity, Q&A on Financial Planning, AI Writing, American Sports

0 views
Manuel Moreale 3 days ago

On values, morals, and doing business

I caught wind of the news that MacStories is back posting on X, the “everything platform”, and that seems to have caused some backlash. I am not a reader of MacStories (because tools are tools and I don’t read about new computers the same way I don’t read about new washing machines) and I’m also not on X. So why do I even care about this news? Well, the short answer is that I don’t. But after the move (I guess because of the backlash?) Federico Viticci, MacStories’ Founder, posted what he called “ An Explanation ” conveniently, not on MacStories’ homepage and not even on the site at all. You should go read the whole thing if you’re curious, but there are a couple of things here and there that caught my attention. The first interesting thing is the attempt to separate Elon and his morals from X. We want to be absolutely clear about something: returning to X does not represent a change in our values, what MacStories or we stand for, or an endorsement of Elon Musk. We abhor Musk’s politics, his rhetoric, his careless disregard for others, and the direction he has taken X. But we also disagree that participating on X means we have aligned ourselves with any of Musk’s views or actions. We reject Musk’s worldview and will continue to treat people with the kindness and respect they deserve wherever they are on the Internet. I’m sorry, but this is not how things work. By being on the platform, by using it to reach the audience that’s on there, you are directly supporting the man. Because the platform only has value because people are on it, and you being there makes the situation worse, not better. The main reason why MacStories is back on X is, surprisingly, money. Boring, yeah I know. Quoting from the post again: It’s not just how we earn a living; it’s one of a small number of independent websites that still cover apps, Apple, and a growing list of topics, including videogame hardware and the automation and productivity side of AI. Readers shouldn’t have to think or care about the business side of MacStories, but we have to, which is why we returned to X. If that is the situation, be transparent. Show how you run the site, how much money the site is making or losing, and make your case. But simply saying “Sorry, running a site is hard so we’re doing a 180 without telling anyone, not even the people who work here” is a shitty move. Especially because the move is pretty obvious: the amount of AI-related content on MacStories has exploded, AI people are on X, and if you want to try to grab some of that money, you have to play that game. Like, I get it, but at least say it and own it. Don’t try to play the running-a-business-is-hard card. Have the guts to put up a post on your precious website where you explain what you’re doing and why you’re doing it. Tell people you want to chase more AI content, because you think you’re now more of a “builder” . You’re an Apple fanboy; have some courage. People with strong values and morals seem to become rarer and rarer these days, especially when money is involved, and it’s so sad to see. And this whole AI moment seems to be exacerbating that. Thank you for keeping RSS alive. You're awesome. Connect via email :: Sign my guestbook :: Support for 1$/month

0 views
Unsung 3 days ago

Movie review: General Magic

★★☆☆☆ 2018, 92 minutes General Magic was a company started in 1990 by some of ex-Apple staffers, working on a pocket communication device and its operating system (Magic Cap, sporting a fascinating room user interface). The product launched in 1994, but swiftly failed in the market, similarly to and concurrently with its chief competitor, Apple Newton . Many alumni of the company – including Megan Smith, Kevin Lynch, Tony Fadell, and Pierre Omidyar – ended up having successful careers in tech after that, to a point that General Magic is referred to as “ Fairchild Semiconductor of the 1990s .” = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/movie-review-general-magic/1.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/movie-review-general-magic/1.1600w.avif" type="image/avif"> The eponymous documentary , released in 2018, was disappointing to me. Sure, it is well-edited, and particularly shines owing to a lot of archival footage shot during General Magic’s happy years. Unfortunately, we get to see none of the sort of details I was hoping to see: no pre-release interface or hardware, no discussion of product nuances, no solid reflection on the company, the culture, or the zeitgeist. The movie drops so many fascinating questions and avenues, but doesn’t really follow up on them: Unfortunately, I kept thinking of the movie as “generic magic” – a sort of interchangeable valorization of Silicon Valley’s “failure is secretly success in the fullness of time,” defaulting to romantic or rousing music, that must have felt obsolete even in 2018. It felt like the documentary could very well talk about one of the many other companies and efforts, which is frustrating, because there was something special and unique about General Magic. I also kept remembering The Soul of the New Machine , Andy Hertzfeld’s own Folklore.org , and even Halt and Catch Fire , all of which more adeptly interspersed personal and emotional drama with specific details you could learn from and take home with you. I think General Magic deserved more. (I do want to recognize that perhaps this wasn’t a movie for me. The documentary is free on YouTube if you want to check it out – it’s consistent throughout, so if you like the early minutes you might like the whole thing.) = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/movie-review-general-magic/2.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/movie-review-general-magic/2.1600w.avif" type="image/avif"> (To play with Magic Cap, go to Infinite Mac , click on Macintosh Garden tab at the bottom, search for “magic cap”, click on magic_cap_simulator.sit, go back to the emulator, double click on the Outside World icon, double click on Downloads, double click on the .sit file, wait for it to unpack, close the Downloads window, reopen it, double click on the Magic Cap Simulator icon.) #apple #history #movie review #review At some point, Andy Hertzfeld reflects on setting a bad example by focusing on small playful details instead of rallying the team to ship stuff. Where did this realization come from and did that change him as a person going forward? In hindsight, did the company benefit or suffer from the culture of freewheeling “superstars” who are also perfectionists? In a strange, brief vignette, a few people talk about not wanting to have any managers – but that is not picked up again or resolved in any way. How did the company manage to hire all of this past and future talent? Are there any repeatable lessons in here? During the launch, there is a brief slide with “whole person thinking,” which felt unique for a tech product reveal. Near the end of the movie, Kara Swisher talks about worrying how mobile devices can affect and perhaps even impair human-to-human communication. Those two threads are not connected. About the only lesson learned and spoken out loud comes from Tony Fadell, who says: “with the iPod, we iterated a lot faster.” Would this have mattered if the premise of the movie seems to be that General Magic was too early on the market anyway? There was a mention of a skeuomorphic room interface done pretty much overnight by Andy Hertzfeld. From the perspective of today, it is the device’s perhaps most distinctive characteristic, but it’s not covered more than that one mention. In another very rare specific example, there is a beat talking how iPhone’s (very controversial, early on) software keyboard owes its existence to General Magic’s team trying that first. How did whoever created it feel about it? An external observer suggests that General Magic missed the early ascendance of the web, but we don’t see anyone from the company reflecting on it.

0 views