Latest Posts (20 found)
Unsung Today

“A quick internet search should provide you with step-by-step instructions.”

In 2018, Tom Cruise and Christopher McQuarrie teamed up to make a short PSA about the dangers of frame rate interpolation, a.k.a. the soap opera effect : = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/a-quick-internet-search-should-provide-you-with-step-by-step-instructions/yt1-play.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/a-quick-internet-search-should-provide-you-with-step-by-step-instructions/yt1-play.1600w.avif" type="image/avif"> It’s a strangely boring video from the men known for exciting filmmaking, but it captures a fascinating debate. TL; DR of their argument: Home TVs have an option to take a typical 24fps movie and create interim frames to make it feel like a 60fps production. It’s often on by default, but you should turn it off. What’s interesting is that 60 or more fps is indeed in some ways objectively better: you see smoother movement, and you can notice more. Once you start paying attention, any whip pan or drastic movement in the cinematic 24fps appears very choppy… …and, of course, that is completely missing the point. The argument for 24fps is that it’s simply cinema’s vernacular, right next to anamorphic lenses with their blue lens flares and vertical stretching. The movies are not about conveying information, but they are about conveying a certain feel . (And also, about tradition.) I’m mentioning this because I spotted a similar battle happening when it comes to scrolling. Go to any page in Vivaldi and hold a down arrow key for a while. That scrolls through the page, moving it in chunks, reacting immediately to the initial key press and then the synthesized autorepeat presses: [force height limit] Firefox listens to the keys the same way, but it “upgrades” the choppiness to a 60fps movement by interpolating it: [force height limit] And Safari does something different altogether – it approaches it more like a game physics engine would. It only listens to ↓ down (when it turns on the scrolling “motor”) and then ↓ up (when it turns it off): [force height limit] This has an interesting effect: the movement is smoother than even Firefox’s – that’s because it doesn’t try to straighten something choppy, but it is itself smooth, by nature. At the same time, this approach feels a bit loose, like driving an old 1960s car. It also takes away any control of speed; Safari pages always scroll at this rate, no matter your keyboard settings. (Arguably, however, control via the key repeat rate that other browsers respect is also an illusion, as it affects typing as well. Would you ever change it just to control the scroll speed?) Of course, the analogy to frame interpolation doesn’t really make sense. Safari is not a soap opera, Firefox is not a smooth motion effect on your TV, and Vivaldi is not a cinematic experience. In contrast with movies and television, I don’t think there are any expectations or tradition here. In this particular context, I believe Firefox and particularly Safari are better because this is about conveying information, and smoother scrolling does help your eyes and your brain connect all the little befores with all the little afters. But I am sharing it mostly as a reminder that a keyboard is really just a button board by a different name. And sometimes it’s good to look at a key hold, and decide: #interface design #keyboard #motion design #youtube is this a sequence of pulsating key presses, or is it a button being held down and then released up?

0 views

How to see hidden East Anglia as a US service member

River Lark and old mill buildings, Mildenhall  by Bikeboy Today I saw a question posted on Reddit from someone who’s deploying to Mildenhall from the US. They wanted insights on how to dive deeper into English culture as a visitor, and I since I grew up nearby and I’m waiting on training I ended up writing a long answer . I’ve no idea if search engines even work any more, but in the spirit of leaving breadcrumbs for anyone with similar questions, I wanted to share a version here too. As someone with family still in Mildenhall and who grew up nearby, but who has lived in the US for 25 years, here are some thoughts for whatever they’re worth. The area has hosted Americans for over 80 years, so the locals are very used to their presence. I even had some USAF kids in my school classes, from families who wanted them to mix outside of the base. I still remember the day one brought in his dad’s pressure suit from a Blackbird for our equivalent of show and tell. I don’t know whether going to an English school was something they benefited from, but I believe it’s still an option if you’re interested. For us it was great to have friends who could bring us American candy. The house rental market is very much set up for visiting military, for better or worse. You’ll have no problem renting, the landlords know they can get help from the base staff if there are any issues, but you’ll also find most places are set up for fairly short term deployments of a couple of years or so. This means you’re not likely to get a rose-covered cottage, but something more functional. You’ll also find it hard to get away from other Americans without traveling a few miles. Photo of the Norfolk Broads by Russell Smith Don’t let any preconceptions of the English countryside mislead you. I find the Fens beautiful, but they’re basically drained swamps, very flat and agricultural. I highly recommend experiencing it from the water, with a narrowboat tour through the Norfolk Broads. You can actually travel pretty much the entire country by connected canals, and some people live on boats year-round. It’s a blue-collar rural area with all the same drugs, violence and bigotry you find in similar areas in the US, though maybe not as obvious. I’ve lost too many friends to suicide and drunk driving, and ask a local about Roma people if you want to hear some ugly attitudes. Different pubs are focused on locals, tourists, and service members. You’ll probably find the locals’ pubs give you a bit of a cold shoulder at first, but that’s also where you’ll get the most authentic interactions if you stick it out. You should also check out the local churches. Even if you’re not religious, you’ll be welcome at a service, and you’ll find them amazing places to visit. I’m an atheist, but Over (confusing village name) St Mary’s is still a special place for me, with a documented history back into the 1000’s, and parts that likely predate that. It’s also been described as having the best gargoyles west of Notre Dame! Churches generally remain unlocked during the day, and if not, the graveyards can be beautiful and fascinating in themselves. There are often local history guides available for purchase (unattended on the honor system) inside the churches. On a similar note, search for local archaeological sites, there are often visitor days or even chances to volunteer. Maybe you’ll find a bog body or a hoard of treasure ? West Front, Abbey Gardens – Bury St Edmunds, by Jim Linwood There are also a lot of small towns with amazing history and architecture, like Bury St Edmunds and St Ives, that don’t have the crush of tourists you’ll find in Cambridge. Look up market days too, it’s often a lot of tat but mixed in there can be some gems. They tend to be aimed at bargain shoppers too, not bougie like farmers markets here. St Ives’ market is over 900 years old , and still held weekly on Mondays, with special events on Bank Holidays. There’s a weird love/hate relationship with the US. I had friends who insisted Americans aren’t funny, while being addicted to shows like the Simpsons (I’m dating myself I know). There’s also strange holdovers like believing there’s still a “special relationship”, the US is “over the pond”, and Americans are our “cousins”. You’ll likely get a lot of questions (or digs thinly disguised as questions) based on stereotypes, news headlines, and people who think their holidays in Orlando give them in-depth knowledge of the US. There’s still a lot of genuine affection for America though. World War II looms large in the British psyche, with war memorials listing the dead in most villages, and we know it would have gone very differently without American help. Bury St Edmund’s War Memorial I wouldn’t worry too much about American habits, but it can be hard to gauge British levels of friendship as an outsider. You might hear “we should get a beer sometime” and try to set a date, only to discover that was just a polite way of saying you’re on good terms with them, like “we should get lunch”. I’ve also had British friends complain that Americans behaved in a way that was perceived as being very close, but then failed to stay in touch with Christmas cards once they left. I feel like the US is a lot more set up for people moving a fair bit, and Britain has a lot more people who live where they were born, though part of that may be because I’ve always lived in big cities here. On the surface a lot will feel familiar, especially because of the shared language, but with a bit of patience and exploration off the beaten track, you will start to see how foreign and complex the UK actually is. I think you’re approaching this with a great attitude and hope you and your family have a good time.

0 views
Unsung Yesterday

YouTube’s right click relay

An interesting mechanism in YouTube I just learned of. If you right click once, you get their context menu: = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/youtubes-right-click-relay/1.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/youtubes-right-click-relay/1.1600w.avif" type="image/avif"> If you right click again , you get the native browser menu: = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/youtubes-right-click-relay/2.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/youtubes-right-click-relay/2.1600w.avif" type="image/avif"> This is clunky and not discoverable, but I can see YouTube team’s bind here. As far as I know, it is not possible to extend browser’s native right click menu (even with user’s consent), or invoke it in some other way (so that, for example, they could have an entry point to the native menu in their menu). At the same time it does feel like the correct use of a right click menu, to host quick functions like Miniplayer or Copy Video URL At Current Time or Copy Embed Code – and, you can also see how Chrome’s own menu has a lot of cruft in it. In a way, this might be what is often called a “progressive enhancement.” It is better than blocking the native menu altogether, and I cannot think of any smarter alternative given the constraints. But it is clunky, as things often are when websites venture out to become web apps. #mouse #web

0 views

Concurrent Servers: Part 8 - Go

This is part 8 in a series of posts on writing concurrent network servers. In this part, we'll switch to Go and see how it tackles the challenges described earlier in the series. All posts in the series: This post assumes a basic familiarity with the Go programming language. As before, we'll start with a sequential server for the basic state machine protocol presented in part 1 . This is the main function: As in the previous parts, the server is "infinite"; it never stops serving new connections until it's explicitly killed. This is the function implementing the protocol for a single client; it takes a net.Conn value that represents a socket with a client connected on the other end: Rather than directly exposing OS threads, the Go runtime implements its own M:N scheduling of lightweight goroutines on top of OS threads. Using goroutines in Go is cheap - both in terms of syntax and developer effort, and in terms of system resources . Here's a version of our serial protocol server that serves clients concurrently by launching a goroutine for each client. The part of the code that's different from the previous sample is highlighted: The concurrent modification in this case is particularly simple because the server is infinite; there's no point waiting for these goroutines to finish (and hence no need for a sync.WaitGroup ). The parameters for server.ServeSerialProtocol are lexically captured from the enclosing scope and its return value is handled by the surrounding closure. Because goroutines are very cheap, this server is very unlikely to run out of resources due to launching too many goroutines; in fact, it will probably run out of something else - like file descriptors for sockets - first. However, sometimes it's still useful to limit the degree of concurrency - even in Go, and we'll discuss some approaches to do so in the following sections. Here are some scenarios in which it makes sense to limit the degree of concurrency in Go programs, even though goroutines are cheap to launch and operate: Let's switch to the primality testing server from part 4 for the rest of the post, because it represents a somewhat more realistic workload. As a reminder: the server receives numbers, simulates blocking by sleeping, and returns "prime" or "composite". The unbounded one-goroutine-per-client version looks almost identical to the previous code sample, except that the goroutine invocation calls another function: Where ServePrimeProtocol is [1] : The simplest way to limit concurrency in Go is by using a counting semaphore, implemented with a channel: The channel sem serves as a semaphore; note that it's a bounded channel with a maximal size. A token is acquired by sending to the channel, and released by receiving from the channel. When the channel is full, the send operation sem <- struct{}{} blocks until a token was removed by some other goroutine [2] . The type of the channel is struct{} which means "empty", or "no data". This is idiomatic in Go for channels that are used solely for their semantics, not to send/receive any actual data. Since launching goroutines is cheap and limiting concurrency is easy as shown above, the "worker pool" pattern is often unnecessary for scenarios like our server. Still, it has occasional uses (such as when workers have to maintain some non-trivial state across tasks) so it's worth discussing it here. Here's a variant of our primality testing server that uses a worker pool: A fixed number of worker goroutines is launched; these goroutines all receive "jobs" from the same channel. In the Accept loop, each client connection is sent to this channel as a new job and is picked up by the next available worker. As mentioned before, you would typically see a sync.WaitGroup somewhere to ensure clean shutdown of goroutines, but in our case it isn't necessary because we have a server that never exits. Do programmers have to resort to async / event-driven programming in Go? In my experience, almost never. Go was designed from the bottom up to be suitable for large-scale concurrency; goroutines are very cheap to create, have a tiny memory footprint and switching happens very quickly, all in user space. Measurements I ran back in 2018 have shown switching times of ~170 ns, as compared to 1-2 microseconds for threads on Linux. Moreover, Go already uses event-driven loops like epoll underneath for I/O. Goroutines that wait for I/O like sockets are effectively "parked" and consume no resources (beyond their small memory footprint); they are woken up by Go's runtime when their I/O descriptors are ready - this is very similar to how asynchronous programming works! That said, some people certainly do try to stretch their resources even more with direct asynchronous programming in Go when millions of streams are handled concurrently. All I'll say is that this is very rare, and an overwhelming majority of users never have to do this. In 2018, I wrote a post named Go hits the concurrency nail right on the head , and after several more years of active coding, I fully stand behind that statement. Go is extremely powerful and ergonomic for concurrent programs; while other environments go to great lengths to implement async-await style event loops in libraries, in Go it's already baked into the core language and runtime. You want event-driven I/O with very lightweight green threads that can also execute blocking tasks without worrying about the function coloring problem ? Go has you covered. All the code for this post is available on GitHub . Careful readers will note two issues with this code: (1) the protocol assumes the complete number is read from the socket in a single conn.Read call, and there's no framing - separation between distinct numbers; (2) the prime checking loop uses i*i which may overflow for large numbers. These issues are consistent across all the versions of the prime server in C, Python, JavaScript and Rust in earlier parts, because my focus was on the simplest possible code to demonstrate a point about concurrency. Part 1 - Introduction Part 2 - Threads Part 3 - Event-driven Part 4 - libuv Part 5 - Redis case study Part 6 - Callbacks, Promises and async/await Part 7 - Rust Part 8 - Go (this part) Tasks may be compute intensive, and the CPU capacity of any server is inherently limited. If too many concurrent goroutines compete for limited CPUs, they will all make very little progress. It may make more sense to have fewer tasks that complete in a reasonable time. Protecting potentially limited downstream resources, such as concurrent DB connections or other services. For example, if the server has to send requests to other services for each task, and these are rate-limited, concurrency will have to be carefully managed. Security reasons when work is dictated by clients; malicious clients can overload and crash a service that's too eager to serve, making it unavailable for legitimate clients.

0 views

A Syncthing and SQLite Gotcha

So, I have this little app, Epoch , that I use to keep a journal. It’s a tiny Rust web app that runs as a systemd service and uses SQLite as the database. I use a desktop and a laptop regularly, and use Syncthing to synchronize them, including Epoch’s database. That way I can use the app on both devices without needing a server to synchronize them, the tradeoff being that I have to make sure the sync is finished before performing any mutations. But I had this bug. Say I edit today’s entry on the laptop, come home, wait for Syncthing to finish, then I’d open today’s entry on the desktop, and the text would be missing. It’s not that the server is holding a lock on the file and preventing the sync: opening the database with the command line tool shows the new text is there. If I restart the server, Epoch can read the new text. My mental model was: The rusqlite object points to the database file. Syncthing swaps the file’s contents from under it. Subsequent queries go to the new file. Turns out there’s a very important part of POSIX filesystem semantics I was ignorant of. The standard way to replace a file safely (i.e. atomically) is the system call: Which Syncthing uses. This I know. What I didn’t know is: what happens if other processes had open file descriptors pointing to ? Do they see the new contents? No: those processes can keep reading and writing to the old file object , but the file is orphaned in that no path points to it. And once all file descriptors are released, the file becomes inaccessible. I’m used to thinking of filesystem operations in terms of “this syscall takes a path and gives you a pointer to the file, which you mutate directly”. Whereas works at the level of directory entries: it atomically mutates the mapping from pathnames to files but doesn’t touch files at all.

0 views
Ahead of AI Yesterday

How Claude Watermarks AI-Generated Text

I recently posted a Substack note about Claude’s new watermarking process and implementation. Since it’s such a popular topic and sparked such a lively discussion, I thought it might be interesting to go into a bit more detail when explaining how it works. Instead of the usual text article, I recorded a little lecture on the topic (to change it up a bit from my usual articles). So, below is the video along with a transcript. Originally, I planned to make 10 slides and record a short 10-min video. However, while putting it together, I added some crucial details here and there, resulting in >50 slides and a 48 min recording. I hope that this now explains it well, though! Happy watching! I also have a YouTube version if you prefer using the YouTube player And here is a link to the slides Subscribe now Note: The transcript below is slightly edited and cleaned up for readability but preserves the overall order and flow of the video lecture above. Slide 2 of 52, time stamp 0:00 Hi everyone. So, a few days ago, Anthropic announced that they will watermark the text outputs of their Claude models. I then did a social media post briefly explaining how that works. And yeah, this was quite the popular post. So not the watermarking itself was popular, but I guess the explanation or the mechanism behind it. Then, it might be worthwhile expanding this a bit to explain it in more detail, because this post only had one figure, and there were a lot of questions and discussions. So, I thought, well, let’s make a few more figures. I actually originally planned to do like 10 slides and walk you through it. It ended up being 50 slides, but I hope this really explains how this watermarking technique works well, how watermarking itself can fail or be removed, and so forth. So I think it might be an interesting topic because a lot of people use LLMs these days and also consume a lot of text on the Internet that might be generated by LLMs. And now there’s going to be this watermarking, and there’s this, I guess, fear of watermarking making text worse, or what’s actually the benefit of this watermarking? And so what does it mean? And I think if we understand a bit better what watermarking is, that goes a long way, and then we can make up our own minds about whether that’s a good thing or not, and so forth, like the pros and cons. So, my goal here is really to explain how the underlying mechanism works and how they are going to implement this type of watermarking, text watermarking. Slide 2 of 52, time stamp 1:41 It’s also a great example to illustrate why understanding things from scratch is actually quite useful. This watermarking technique is also a nice way to explain how conventional models or LLMs in general work under the hood. So yeah, you may know I like doing things from scratch. Like, I have my books: Build a Large Language Model From Scratch, Build a Reasoning Model From Scratch. I have some articles labeled from scratch. So, for me, “ from scratch often includes coding. So this one will not be coding-related, but coding from scratch is actually a very, very useful technique because it really helps you understand how something is implemented. And then from that we can derive our understanding, figures, concepts, because if we don’t really implement things, if there’s no code, it’s really sometimes ambiguous. And of course, you know, as I realized, not everyone is coding from scratch anymore. Like back in the day, coding something from scratch was all we had. I mean, there were only humans coding. Nowadays, coding can be done by LLMs. However, that doesn’t mean reading code is no longer useful, because it carries a lot of information. So in this case here with this watermarking, spending some time coding an LLM from scratch really makes you realize how this sampling inside is implemented. We still have some relevant code snippets. And then that really, in turn, helps us understand, oh, the watermarking is applied at this position, and this has so-and-so consequences and so forth. So I think even though people may not be coding from scratch, at least not all the time anymore, it is still useful being able to, let’s say, build something from scratch for educational purposes to understand something deeply and then also for research purposes to manipulate this in a transparent way that is not hidden away in tons of layers of abstraction. But that aside, I think it’s just a coincidental nice relationship here because, for this slide deck, I actually used a lot of figures from my from-scratch coding materials. Slide 3 of 52, time stamp 3:58 So a few days ago (this is August 14), there was this article, How Claude’s Text Watermark Works, and there was this article here; it’s just like a screen recording, so it can have everything in the slides, but there’s plenty of detail. They updated it actually a couple of times, so originally when I read this, it was a way shorter. Still, it is very, I guess, conceptual; there’s like this overview, and there’s, I mean, there’s not a single figure in there. And so it’s kind of still hard to understand what they’re trying to do. So they explain a lot about why they’re going to do it, but they don’t explain how. They’re linking to one paper somewhere there, which is very technical also. So I do think it makes sense maybe to take a step back and start at the beginning to kind of understand what they’re trying to implement here with this watermarking technique. And so the motivation, by the way, of watermarking is for them to identify if someone posts some text that they can say, oh, this text was generated by our Claude Opus 4.8 model, for example, so that they have a way to tell, OK, this text is AI-generated because it carries this watermark. And this watermark is invisible to users, so only they can decode it and find out whether the text has their watermark. Why can only they do it? We will get to that later in this (hopefully not too long a video), but one thing at a time. Slide 4 of 52, time stamp 5:35 So I wanted to start with a brief prelude to explain how text generation works in LLMs, because based on that we can then more easily understand how the watermarking works and that this is actually not a huge, expensive thing on top of it. It’s really just like a minor, I guess, tweak inside the regular text generation process. Slide 5 of 52, time stamp 6:01 So when we are using something like ChatGPT, for example, let’s say I ask the question, the capital of Germany is, and yeah, ChatGPT or other LLMs, so this is just like an example would, for example, answer “Berlin”. So here, in this case, it’s generating two tokens, like “Berlin” and the period. But for simplicity, let’s assume it’s generating one token. So the next token is the “Berlin” token. How is this token generated internally? What is happening under the hood when we type something here like the capital of Germany is and receive a token like “Berlin” back? What is actually going on there behind the scenes? Slide 6 of 52, time stamp 6:41 So in the next couple of slides, I want to briefly talk about what happens under the hood when this next token is generated. Slide 7 of 52, time stamp 6:50 So assume again that our prompt is the capital of Germany is. And the first step here is to convert this into token IDs. So tokenizing it and converting it into token IDs is one of the main steps at the beginning. This is outside. It’s not inside the LLM; it’s outside of the LLM. So we are simply converting the text into token IDs. It’s just a format that embedding layers can work with. Slide 8 of 52, time stamp 7:22 And then this passes through the LLM. And the LLM gives us a score distribution for the next token. Slide 9 of 52, time stamp 7:31 So again, this is just like a brief overview of how LLMs work internally. So I’m not covering the LLM machinery itself. I talked about it many times in my other From Scratch LLMs videos and books. The important part is that when we generate the next token (for example, “Berlin”), we have, at this point, a distribution of scores. So this is the output produced by the LLM. Here in this case, we’re looking at logit values. So these are just scores from minus infinity to plus infinity, like a range of scores. Here’s an example, ranging from about -8 or -9 to 20. We could convert these into a probability distribution, but technically, it’s not strictly necessary depending on how we sample. But so you can think of the logit values as the raw scores. And the raw scores go over the entire vocabulary. Slide 10 of 52, time stamp 8:39 That means every possible word that the LLM could generate. Now here, in the vocabulary index, a certain value (index position 19,846) receives the highest score. So I spread out the distribution. If you would run this prompt through an LLM, you would even see something more extreme: that everything is, like, very, very, very close to zero. And “Berlin” would probably be much, much higher even. But just to show you a few, you know, like peaks here so it looks a bit more interesting, I kind of zoomed in; in and spread out the distribution a bit. Now here, “Berlin” is the highest score because you can think of it as the most, I guess, probable or plausible next token if I have a very specific prompt like this. So the other ones, I mean, it could be something like Hamburg or Munich that the LLM might guess incorrectly. But nowadays an LLM should be fairly certain that “Berlin” is the correct answer here. You are also seeing here the vocabulary index. So that’s like over the whole vocabulary. Nowadays, LLMs have like 250,000 possible tokens as output. I’m truncating it here from 19,800 to 19,900 because there’s just so much space here on this slide. If I would have a very realistic vocabulary of 250,000 words, everything would be so narrow that we would barely even be able to tell or see anything on this distribution. So this is just truncated for educational purposes. The important point is that in regular text generation, we get this score distribution. Now, what we do is look at the highest score. Slide 11 of 52, time stamp 10:33 I will get into more detail later on how this is selected. So it’s not necessarily precisely the highest one, but for simplicity, assume we are taking the highest score here. And in this case, it’s 19,846. Slide 12 of 52, time stamp 10:52 And this score is then detokenized, and we get “Berlin” back. So that is the process here on this slide: from an input prompt to conversion into token IDs and tokenization, passing it to the LLM, getting this score distribution, getting the next token, and converting it back into text. Slide 13 of 52, time stamp 11:12 And then this text is appended to the input. So if we have a question that requires multiple output tokens, we keep going in this loop until the answer is complete. That usually means that the LLM generates an end-of-text token, for example, here. For simplicity, I’m showing you only one iteration where it generates one token. But yeah, as I said, it would kind of continue like that, where we are feeding back the modified input to the LLM for the next round. Now, how do we actually sample this next token here? Slide 14 of 52, time stamp 11:44 I briefly said, well, we could just technically select the highest one, the one with the highest score. This is called greedy decoding. That’s one way to do it. But most LLMs, like if you use them, they don’t do greedy decoding where they always pick the highest one. Because if you ask it on some other prompt, it might not be what we want to always have the highest score, because then it would memorize the training data. It would always kind of give the same response and so forth. So we actually often want some variation in the outputs, but not so much that it generates random stuff. So how it works is that, when we sample here from this distribution, we first typically convert it into probability scores. Slide 15 of 52, time stamp 12:28 So here I just have these scores shown in this plot. I’m just using NumPy for simplicity; whatever tool you use (e.g., PyTorch), the same concepts apply. But let’s assume we have the scores here in NumPy. So what I would do is I would compute the softmax. Technically, I would use a softmax function implemented in Torch or PyTorch, for example, that is numerically stable for both large and small values, including very high positive values, very low positive values, and very high negative values. Here I’m just writing it out like that. That’s the canonical softmax, just to make it a bit more readable. But the details don’t matter here. Slide 16 of 52, time stamp 13:20 What matters is that after this conversion, the scores here, I mean, there’s only so much space on the slide, but the scores here, they would add up to one. So it’s essentially like a renormalization. So they would be normalized to sum up to one. That’s all that the probability conversion does: the softmax conversion. So then once we have these probabilities, we can use a random number or, like, a random sampling algorithm. For example, here in NumPy, we could use the choice function or method. So this is with a specific random seed we are passing to the vocabulary indices. And then, and that’s the important part, we are passing the probabilities as the weights. So, these, essentially, yeah, are like: “How likely is a certain token to be selected?” So, for example, if “Berlin”, after this normalization step, the softmax step, has a 99% probability and the other ones together have a 1% probability, then if we would sample 100 times, 99 of the times, we would get “Berlin”. In realistic LLMs, for example, that are well trained, “Berlin” might receive a probability of 99.999999 or something like that. So you’re almost certainly always sampling “Berlin” because it’s very confident that the answer is “Berlin” in this particular case. So yeah, that is how we would sample from this distribution. There are modifications like top-k sampling or top-p sampling where, let’s say, just for simplicity in top-k sampling, we would select the top 100 tokens and then apply this random choice only to the top 100, the 100 highest-scoring ones, so that we don’t get nonsense tokens in there. For this example, it doesn’t really matter. I mean, it’s just like another thing to explain, so I’m skimming over this. So you can maybe assume that this is already the top 100 tokens using top-k or something like that. Slide 17 of 52, time stamp 15:37 And so, for example, here’s an example. If we sample 10,000 times with a probability of “Berlin” being very high, 99.9, we would sample “Berlin” 9,997 times, sample the word “Hal” twice, and one “Moh”. And these are basically nonsense tokens. It rarely happens that, in this case, the LLM might produce nonsense because, as I mentioned before, I spread out this distribution a bit to make it more interesting. A real LLM would probably, 10,000 out of 10,000 times, sample “Berlin” because the probability of “Berlin” is so high. But this is for illustration purposes. Slide 18 of 52, time stamp 16:21 Now we briefly talked about how LLMs work under the hood, which I think is kind of an interesting concept in itself. But I’ve talked about this many times before, so I don’t want to bore you. I just wanted to set up some context for now, explaining how this watermarking works. Slide 19 of 52, time stamp 16:41 So, we mentioned that we select the highest-scoring token when sampling. Or we use this probability sampling, which will lead to one of the highest-scoring tokens being selected most of the time. Now here’s another example without watermarking. I changed the prompt slightly. Now the prompt is: today’s weather is “cold,” and a possible answer could be, for example, “gray” or “overcast”. So in contrast to the “Berlin” example, I would say “gray” and “overcast” kind of are interchangeable. They are both reasonable next tokens for this prompt, given the goal of completing this text or writing the next token. So it’s almost like a coin flip which one we want to select. There is not really an objectively worse one of one or the other. So when we do the random sampling, because they also have relatively high scores and their scores are similarly high since they are both plausible tokens, we might get one or the other. So almost half of the time we would get “overcast”, and almost half of the time we would get “gray” if we repeat the sampling multiple times. And that’s how LLMs often end up with different answers if you provide the same prompt. If you use the same prompt and you ask the LLM multiple times, you often get slightly different answers. And that’s because at certain positions, two possible tokens are almost equally likely, so it will choose one or the other. And that token would then influence all subsequent tokens, and so forth. Slide 20 of 52, time stamp 18:25 Now, I wanted to briefly talk about random number generation. So, for example, if we use a random number generator like this, it will generate a random sequence of numbers. If I run it again, the sequence of numbers is different here. So you can see every time we produce five numbers, they are different. If I set the random seed here, like one, two, three, and I run this multiple times, we still get random numbers, but they are now all the same, right? So they are still random. If we use a random seed, we still get random numbers that are different from each other, but they are reproducible. So whether we use a random seed or not, we still get random numbers. But with a random seed, we get a reproducible sequence of numbers. So keep this in mind: this is just like a little primer, and we will use this concept in a few moments. Slide 21 of 52, time stamp 19:26 So, for example, I mentioned before that we might get either “gray” or “overcast” if we randomly sample. Now, if we use a specific random seed like 42, we would always, for example, select “overcast”. I mean, it’s still a random selection, but we make it deterministic. In this case, given this prompt, the model will always select “overcast”. Slide 22 of 52, time stamp 19:49 If we use a different random seed, the model might select “gray”. Every time we sample, it will always select “gray” as the next token. So it’s still random sampling, but we are making it deterministic based on the random seed. Slide 23 of 52, time stamp 20:04 So, in watermarking, Claude watermarking is kind of like the idea that it sets a random seed. But this random seed, instead of being like a number that is fixed based on, I don’t know, someone writing down a fixed number, they’re using a secret key that is essentially like an API key, a secret key, and from that key, together with the four previous words, they derive this random seed essentially. But the idea is that if I go back one slide, it’s the same as here: there’s essentially a fixed random seed, and that random seed always selects the same next token. Okay, so instead of using random seed 99 here, for example, they have a secret key and also use information about the previous tokens to derive this random seed. But more on that later. Slide 24 of 52, time stamp 21:06 So the idea is that watermarking makes the text generation more deterministic in certain positions. So, for example, if we have these plausible texts on the left side. So if I have a text that says, > The weather today is cold and I may either pick “overcast” or “gray”. And the next sentence could be, > and then “light” or “gentle” They’re both interchangeable again. > And then breeze is “moving” or “blowing” through the trees, and the streets seem “quiet” or “still”. Which means basically I could say either “quiet” or “still”. So there are certain positions in the text where we have token choices where they are almost equally likely, like we have seen before. So that means if we are, this is without watermarking, if we are running the prompt, or given the prompt through the LLM, we might sometimes get this answer here, sometimes this answer, and so forth. And based on the number of positions, we might have 128 possible answers here. And of course, the longer the text, the more positions we have where we can have terms interchangeably, the more combinations, or the more output texts, there are. So, for example, again, one possible output text could be > The weather today is cold and overcast. A light breeze is moving through the trees, and the streets seem quiet. I think I’ll stay home and read a book with a cup of tea. So that is one possible text. Another possible text is > The weather today is cold and gray. A gentle breeze is blowing through the trees, and the streets seem still. I think I’ll stay inside and read a novel with a mug of tea. By the way, it’s also actually raining outside. I don’t know how good this microphone is, but it’s kind of a very fitting context here. But yeah, the bottom line is that you can see there are two very reasonable texts here being generated, and there are more combinations. So they are all reasonable. There isn’t one that is necessarily better than the other. They’re just, you know, slight variations. And if we don’t use watermarking, we might get either one, or it’s just random, right? Because of the random sampling, we might get one or the other. Slide 25 of 52, time stamp 23:31 Now, if we fix the random seed, as I mentioned before, for example, if the random seed is 99, we might always get this text here. So, using a random seed, we can kind of fix which answer we get, because then the random sampling is still random, but it’s deterministic in the sense that it’s reproducible. It’s always going to be the same then. Okay, so that is still without watermarking, now with a random seed. Slide 26 of 52, time stamp 23:59 And the watermarking is essentially doing the same thing. Now, instead of just using a simple random seed, they have a so-called random key, where this random key is involved in selecting the text, essentially. But what we can already say is that, in the Claude blog post, they say the watermarking shouldn’t make the text worse. If we look at this mechanism, yeah, it makes sense why it would not make the text worse. By the way, I’m not defending watermarks here. I’m just trying to explain. So please don’t kill the messenger here. But what I’m trying to say is that the watermarking is nothing else for the end user than fixing a random seed and making this sampling kind of deterministic, if that makes sense. Slide 27 of 52, time stamp 24:47 Okay, so the summary so far is without watermarking. We often sample without a random seed because I know most people don’t even use one. I honestly don’t think you can necessarily do it with the Claude and OpenAI APIs. I know you can do it in Ollama, but I also always had some problems with that because I used Ollama in one of my books for the bonus material to generate some texts. I was fixing the random seed, but it still wasn’t always deterministic, and so forth. So it’s tricky. Your mileage may also vary, depending on the software version and so forth. Anyways, so without watermarking, we have this random sampling. With watermarking on the right-hand side, we still have the random sampling. But in addition to just a random sampling being fully random, we have this watermarking key. And this watermarking key is passed to the random seed generator to set a specific random seed, making this deterministic. But it’s essentially very similar, and like I mentioned, there’s a lot of benefit in terms of understanding things from scratch. And now we know essentially where this watermark is applied to. So this is essentially applied to the sampling. It’s not applied inside the LLM, which is actually cool knowledge. So they don’t need to train a new LLM for that. They can just use an existing LLM, and they just apply it at this sampling stage. They don’t have to retrain anything or anything like that. So yeah, that is actually interesting, right? Slide 28 of 52, time stamp 26:21 But we are not quite done yet. I would also like to talk about how we can understand or see whether text is watermarked. So detecting the watermark is only possible if we have access to the key. So, for example, if we have these different texts, and essentially, after the text was generated, you find some random text on the internet (for example, you find this text number four here on the internet somewhere), you want to know: is this watermarked? Well, it’s impossible to know because, in order to know, you would need the watermarking key. You need this scoring function, and then you have to score basically the text with a scoring function. And then the idea is that if the score is above a certain threshold, then the text is watermarked. Otherwise, it’s not watermarked. But as the end user, we can’t do this because we don’t have this key. So the key is not available to us. Only Anthropic will have the key. However, in this blog post, they mentioned that they are providing it, of course, or they’re going to develop an API for that that they will make available. I don’t know. Honestly, I’m not affiliated. I don’t know the details. I was just reading this in this blog post. That’s all I know. So that API might as well be private for some companies, like, let’s say, X or Substack Notes, when they want to label AI-generated posts. They may make it public for end users to use. Who knows? We will have to wait on that. But yeah, so the bottom line here is that watermark detection is only possible if we have this watermarking key or, of course, the API that they are going to develop. Slide 29 of 52, time stamp 28:03 Now, removing the watermark is interesting. So now that we know how the watermarking works, we also know the shortcomings. I mean, this is really highly dependent on specific tokens in certain positions. So, for example, in this given text, if these colored words or tokens are the watermarking positions, we know that we could remove the watermark by editing this, right? If we change all the words at these positions, we would be 100% able to defeat this watermark. Now, the problem, though, is that we don’t know, right? Slide 30 of 52, time stamp 28:41 So we don’t know where these words are because we haven’t generated the watermark. So we don’t know which positions to look at. So the practical scenario here is that we could just randomly edit the text. So we would randomly change a few words and hope that we change enough positions to edit the watermark. So that would be one way to remove it. And since we also don’t know which are the highest-scoring ones, because that would require us to have access to the LLM and rerun the prompt through the LLM to find out which words are the highest-scoring, we can kind of only guess. So for example, we might say, oh, we replace “overcast” with “cloudy” because we don’t know that “gray” was high-scoring, you know? So in this case, it might be intuitive to say “gray”, but there might be cases where it’s not so intuitive. So what I’m trying to illustrate here is just some general text editing where we are modifying positions, but we are still kind of guessing what a watermark position is. So since we don’t know, we added just a few words here and there. And if we added enough words, that would also defeat the watermark. Slide 31 of 52, time stamp 29:59 So yeah, that was the watermarking in a nutshell. I mentioned that there is a scoring function to find out whether something is watermarked. And I want to do it as a bonus here. It’s already a long video, but as a bonus here, I wanted to briefly also explain how this scoring function works because that is also interesting information. It’s a bit complicated. It’s not essential to understand how the scoring function works. But the reason why they do it the certain way they do is to make the detection cheaper. Because otherwise, if I go back one slide or two slides, if you wanted to check if something is watermarked, if even they wanted to check, they would have to rerun the prompt to get these scores and then apply this watermarking random seed to get this text and then compare. And that would be very expensive because then essentially every text you want to compare, you would have to rerun the LLM. You have to know which LLM, and that would be really unfeasible because you often also don’t even know the prompt, right? So yeah, so they have like a trick that they use to, yeah, I would say, modify the sampling so that you don’t use or don’t need the LLM later on for the scoring stage. And in the blog post, they mentioned that they derived this method from a paper. It was a Nature paper, and this method is called SynthID-Text. So that was like a paper that came out maybe one or two years ago. It was by Google, and they use a similar technique they call Claude watermarking. I don’t know, sorry, I don’t know if they use exactly that technique, but that’s the one they mentioned. Slide 32 of 52, time stamp 31:39 So how does it work? So before we looked at the slides, we looked at the regular, let’s say, overview here, where we have some text. We put it through the LLM. We get this logit distribution and then we sample from the distribution and get the output token. And here, during the sampling, we use the watermarking key and the random seed generator. So this is still correct. This is still what’s going on, but there is a bit more, I guess, nuance to how this token is sampled. So they’re not just using, let’s say, NumPy’s random choice. They’re using something a bit more sophisticated here. Slide 33 of 52, time stamp 32:16 So assume, again, our context is “the weather today is cold,” and we want to generate the next token. So, for example: “gray”, “overcast”, “gloomy”, “cloudy”. “Gray” is 50% probability, “overcast” is 30, “gloomy” is 15. Let’s say “cloudy” is 0.05 and the rest is, let’s say, 0. Here it looks, of course, a bit different. Let’s say that’s “gray” and “overcast”. I’m just reusing this figure. But now imagine these are the most likely ones, like “gray” and “overcast”, and everything else is just very small, except “gloomy” and “cloudy,” maybe. So essentially, think about just a very small vocabulary for this example of four words instead of all these 50 words here, just to make it even simpler. Now, as I mentioned before, we could use ‘sNumPy’s random choice with these probabilities to sample the next token. Slide 34 of 52, time stamp 33:13 And we could use the watermarking key with this random seed generator to make it deterministic and get the certain watermark that we want. But as I mentioned before, this would be very expensive. Not the sampling itself. That doesn’t matter. This is pretty cheap. But the detection later on would be very expensive if we are trying to check random text on the internet. Slide 35 of 52, time stamp 33:35 So instead, what they use, they also use it during the sampling, during the generation, so that it can be reused later during detection. What they use is called tournament sampling. So this is instead of using something like random choice, they use a concept called tournament sampling. And so how does that work? It might look a bit complicated, but it looks really more complicated than it really is, to be honest. So you might have to, I guess, stop the video at some point and just sit with the figure a bit. But I think it is actually simpler than it looks like. It’s like once you get the hang of it, it’s pretty straightforward. But let me try to explain here. So what we have is we have still this context, and then we have these probable or plausible next tokens with these different probabilities. Now they have something they call random watermarking functions. Slide 36 of 52, time stamp 34:35 Here we have three watermarking functions, G1, G2, and G3. In reality, they might have 30, 50, or even more. Here I’m just using three because that is simpler on this slide. It’s just smaller, you know, like it fits better on the slide. Now, if we look at this word “gray”, this might give us a signature 101. With that, I mean, if we use this watermarking key to generate this random seed, and we have three functions, G1, G2, G3. If I put the word “gray”, what I’m skipping here is that usually you put the word “gray” together with the four or three previous words from the context. So it’s “cold” and “gray”. If I put that into G1 together with this watermarking key, I get the value one. Why? Well, that’s just how this function works. It’s like a random function. The random function either returns zero or one. In this case, with this random key and this token, it returns one. With the same key, but a different function, you get the value zero. And then here you get a one again. So if we have more functions (of course, 30 functions), this will be a very long string of ones and zeros. Slide 37 of 52, time stamp 36:00 It’s basically like a bit string, like if you have bits of zeros and ones. Okay. So this is for “gray”. So we get the signature 101 through using these watermarking functions. Now we can do the same thing for all the other ones. So we can do it for “gray”. We can do it for “overcast”, “gloomy”, and “cloudy”. So each one has a different signature here. So, for example, “overcast” is zero, one, zero. “Gloomy” has zero, zero, one. “Cloudy” has one, zero, zero. Okay. So we have these bits here now. The next step is a so-called tournament sampling where we just pair them. Slide 38 of 52, time stamp 36:39 Like, you know, like a soccer tournament, the knockout (KO) stages, or like the playoffs in American football, you always have two teams playing against each other. And that’s kind of like the same idea. We have a pair of tokens, and they’re playing against each other, essentially. And the scores, they come from these functions here. So we start with the first function in the first round. So we have “cloudy” and “gray”. So we look up here: “gray” is a one and “cloudy” is a one. Okay. So one and one. “Overcast” and “gray”. So “overcast” is zero, “gray” is one. So we have zero, one. “Gloomy” and “overcast”. So here we have “gloomy” zero, “overcast” zero. So zero, zero. And then we have “gray” and “gray” again, because we are running out. So we don’t have enough of the others. So we have one duplicate. So this is chosen randomly. And so you have one and one here. Now we look at the results. So this is a tie. In the case of a tie, we also select randomly using, you know, the random seed and the watermarking key. So here, “cloudy” survives. And from this one, G1 is, according to G1, “gray” is the winner because it has the one. So “gray” survives. And then here, “overcast” and “ gray “ are a tie, randomly selected, and “gray” also randomly selected. So we have now “cloudy” and “gray” and “overcast” and “gray”. And we play the next round in this tournament. So in this next round, we use G2. So according to G2, “cloudy” has a zero here. “Gray” also has zero. “Overcast” has one. And “gray” also has zero, sorry. And so, the next stage of the tournament again. Slide 39 of 52, time stamp 38:24 So we have a tie. We randomly select “gray”. And here we have “overcast” as the winner. And so we have “gray” versus “overcast” in the final. And then we look again at the scores. So “gray” has a one. “Overcast” is a zero. So “gray” is the winner. And that’s how the token “gray” is sampled. What is the watermarking key doing here? So the watermarking key, if I go back a few slides, is selected for generating these scores using these random watermarking functions. So the watermarking key determines essentially what values we get at these stages. So the watermarking key is still very important. Otherwise, these signatures would look different. Slide 40 of 52, time stamp 39:11 So we now have sampled the next token. And that’s just how this modified sampling procedure works. We could have used NumPy’s `random.choice`. But the shortcoming of that is that if we want to score random text on the internet, we would have to rerun the LLM. With this technique, we don’t. I will show you in a moment. So this technique sounds like really weird and cumbersome, but it has the advantage that we can now score random text more easily without having to rerun the LLM. So it’s essentially just to make the detection easier and cheaper. Slide 41 of 52, time stamp 39:43 Slide 42 of 52, time stamp 39:48 So, for example, if we have a new text. So I’m just using the same text here, but let’s assume it’s new text. So this is after the sampling, when we are scoring. And let’s say we are discovering this text on the internet. And the text is the weather today is cold and “gray”, and we want to know if this is LLM-generated or not. So we would, or Claude/Anthropic would, have the watermarking key and these functions: G1, G2, and G3. And it would put this text through these functions. For the one position here for “gray”, we would get 101, similar to what we got during the generation process. So this is the same as before. And this has, if we add up these bits, two bits, right? One and one here. So it has two bits of information, let’s say, for simplicity. This is just a really simple illustration. But let’s assume we get a score of two here for the “gray” in this position. If we had a different word here, “overcast,” in this position, we would get one if we get “gloomy,” like we also have one, and “cloudy” one. So I’m just summing over each row here, right? So that’s just like a score we would get at each position. And here I’m only looking at the last position. If I would do this at other positions, I would get a different score at different positions. So, for example, let’s assume at the first position I get a two. Here I get a two. For “today”, I get a three. For “is”, I get a two. “Cold”, two. And “gray”, three. So here I’m applying these watermarking functions as I’ve shown on the previous slide. Slide 43 of 52, time stamp 41:33 And I’m just adding up these numbers across the three functions. And the watermarking functions are very cheap. So you can just quickly run them on the whole text and get these scores. And then based on that, I can compute the average bits. So if I just average over all these values here, let’s say I get 2.23. Slide 44 of 52, time stamp 41:55 Now, if I have slightly different text, so here I swapped “today” with “now” and “gray” with “overcast”. These now get a score of one and one. And if I average over this whole string, then I get a 1.71. And so for that, I don’t need an LLM. All I need is the watermarking key, the random seed generator, and these functions, G1, G2, and G3. And that’s all I need. I don’t need the LLM. And I can get this score here. And what they do is apply a threshold. Slide 45 of 52, time stamp 42:27 So, for example, I mean, they don’t use this exact threshold. But for example, we can say if the score is greater than two, then the text is watermarked. If the score is smaller than two, it’s not watermarked. So here, if the score is greater than two, it’s a yes. So yes, this is watermarked. In this case, 1.71 is not greater than two. So this text is not watermarked. Okay. So that’s just the way we can then detect whether random text on the internet is watermarked or not. It’s essentially just applying these watermarking functions and then averaging over the scores and applying a threshold. Okay. Slide 46 of 52, time stamp 43:13 So again, the tournament sampling is mainly to make detection easier and cheaper. We could also use something like NumPy’s random choice with a random seed or to make the sampling deterministic. But then again, it would be hard to score any text on the internet. Slide 47 of 52, time stamp 43:30 So yeah, the summary is still the same, though. The thing that is different between no watermarking and watermarking is that we are controlling this sampling here with the watermarking key. And inside that, we have this tournament sampling. And yeah, as I mentioned before, detecting the watermarks requires the secret key and the watermarking functions G1 to Gn. Slide 48 of 52, time stamp 43:54 And again, removing the watermark, because I think that’s maybe interesting to some people, would ideally involve editing all the positions here. But since we don’t know which positions are watermarked and internally, they choose the positions so that they have equally likely tokens at those positions. And there might be positions where that’s not true. So here, for example, for “trees”, we might not even have an alternative word that is high scoring so they don’t watermark that position. So they only do the watermarking at certain positions essentially. Since we don’t know which positions to kind of defeat or remove the watermark, we would... Slide 49 of 52, time stamp 44:29 ...have to edit several places in the text. So what I think that means for the future of AI-generated text is that this actually... Slide 50 of 52, time stamp 44:36 …might result in worse AI-generated text. So I think if there’s a person who likes to use AI-generated text everywhere on the internet, let’s say there’s a news website that likes to use AI-generated text to write the news, I don’t think watermarking will necessarily stop them from doing that. They will probably still want to generate AI-generated text because that’s part of their workflow. So I think my guess is that they’ll use another model. Slide 51 of 52, time stamp 45:08 They’ll just use a second model to edit the text to get the so-called edited AI-generated text. So it’s complicating the pipeline. Instead of getting the text directly from Claude, it’s now using Claude to generate AI-generated text, passing it through a local model, and then having edited AI-generated text, which is likely not watermarked anymore. So why a local model? I just think a local model because I think all the providers- the proprietary LLMs, not only Claude, but also Google— I mean, Google wrote this paper, right? So I’m thinking that they are also watermarking Gemini text. And I think OpenAI is probably already doing it or will do so as well. I mean, I’m just speculating, but I’m imagining everyone will probably do something like that because there’s like an EU regulation that requires that. And that’s, according to the blog post, apparently why Claude is doing it. Yeah, so I’m thinking local models may not, at least not yet, implement this watermarking. So I think people will just use a local model and then generate edited AI-generated text. And my guess is it will be slightly worse than the original text because for the local model, you might now be using a smaller model. So, I mean, you could also technically just use the local model directly to generate text. But in my view, editing text is simpler than generating text. So for the generation of the text, you might use a very expensive high-end, I don’t know, like the highest, most expensive Claude model for complicated text. And then you use a cheaper local model to make these surgical edits, essentially. That’s probably what’s going to happen. And why worse? So if we look back at this graphic where we just added random positions, you might be just changing words for the sake of changing them. And then it risks making the text worse. So you might still have generated text, but it’s kind of like it’s edited awkwardly. Slide 52 of 52, time stamp 47:20 But anyway, so my goal here was to explain how the watermarking works and not, let’s say, the worldwide ramifications of that. But I hope this kind of behind-the-scenes, under-the-hood look is useful. The watermarking is not as complicated as it might seem, but I think it was still 52 slides, so it was also not super trivial. So I hope you found this little lecture useful. And yeah, until next time, see you then. PS: If you like more explainers in this style, I don’t post videos to YouTube regularly, but I have accumulated over 300 videos over the years, which you can find on my YouTube channel here . I also have a YouTube version if you prefer using the YouTube player And here is a link to the slides

0 views
Evan Hahn Yesterday

Vim's UserGettingBored autocmd

In short: Vim has a joke autocmd called that doesn’t do anything. Vim’s automatic commands feature, usually shortened to “ autocmd ”, lets you run code when various events occur. For example, you could implement an auto-save feature by binding the event to the command. Vim has over 100 events, from “buffer was created” to “file was saved”. But one of them sticks out to me: . Here’s the documentation: : When the user presses the same key 42 times. Just kidding! :-) When I saw this, I was busy doing something else and it completely derailed me. “I must know more,” I thought. Here’s what I found: Unfortunately, it doesn’t do anything. It only exists in the documentation (and some tests). If you try to use it with somethig like , you’ll get a “no such group or event” error. It’s present in Vim , Neovim , and Vim Classic . It was first added by Bram Moolenaar in July 2000 , over a year before Vim 6.0 was released. The original description was, “When the user hits CTRL-C. Just kidding!” And it didn’t do anything back then, so I don’t think it’s ever been real. In August 2001, he added the smiley face to the documentation . It then read, “When the user hits CTRL-C. Just kidding! :-)” Twelve years later, in 2013, the description changed to its current iteration: “When the user presses the same key 42 times. Just kidding! :-)” In 2022, developer Mike Smith created an unofficial plugin inspired by this joke autocmd . If you press the same key 42 times in Insert mode, a picture of Samuel L. Jackson appears. 22 years later, it’s finally real. Unfortunately, it doesn’t do anything. It only exists in the documentation (and some tests). If you try to use it with somethig like , you’ll get a “no such group or event” error. It’s present in Vim , Neovim , and Vim Classic . It was first added by Bram Moolenaar in July 2000 , over a year before Vim 6.0 was released. The original description was, “When the user hits CTRL-C. Just kidding!” And it didn’t do anything back then, so I don’t think it’s ever been real. In August 2001, he added the smiley face to the documentation . It then read, “When the user hits CTRL-C. Just kidding! :-)” Twelve years later, in 2013, the description changed to its current iteration: “When the user presses the same key 42 times. Just kidding! :-)”

0 views
Sean Goedecke Yesterday

You should never be angry at work

I try not to give a lot of prescriptive advice about working in tech companies 1 . There are many ways to be successful, and every company works differently. If you’re shipping projects and your management chain is happy, it doesn’t really matter how you’ve accomplished it. However, there’s one thing that I do think is solid advice: you should never be angry at work . Anger in the workplace is toxic. An angry colleague immediately becomes a new problem to be managed, not a professional helping you manage problems. When someone is visibly angry in a meeting or in Slack, it kills the entire atmosphere: other engineers will often go quiet entirely, not wanting to make the situation worse. If you routinely “get heated” at work, the best-case scenario is that you’re part of a tight-knit team of confident people who aren’t put off by it 2 . No harm, no foul. But the second someone comes onto your team who’s not so confident, or you have to communicate outside of your team, it becomes a big problem. Healthy workplaces route around anger in the same way that networks route around damage. Emotionally unreliable engineers will get left out of conversations that might cause them to blow up. Decision-making will get done around them in backchannels. I’ve seen this become a self-reinforcing cycle: angry engineers aren’t consulted on key decisions, which makes them angrier, which pushes them even further away from the spaces where decisions get made, and so on. You can often find these engineers bitterly complaining that they keep the company together, but nobody ever listens to them. In my experience 3 , this is almost never true. Engineers who are highly effective tend to get listened to — at minimum by their colleagues, and eventually by managers and product managers who want to extract as much value as possible from them. (One reason this is true is that all successful projects involve working with other people, and if nobody listens to you, you can’t do that.) Why do angry engineers believe they’re important? Paradoxically, anger can be really useful to a software engineer . Angry engineers are rarely the ones holding the company together, but they’re also rarely useless . One surprising thing about working for big tech companies is that some engineers are not just unproductive, but actively net-negative : either because they’re incapable of doing useful work on their own, or because they’re sloppy enough that they create more work than they do, or because they’re so checked out that they literally do nothing. Angry engineers might be net-negative in a cultural sense, but in terms of literally solving tickets and shipping features, they’re usually well above average. Why is this? Anger often comes from caring about your work, and caring a lot is sufficient to make you a competent engineer . I’ve never worked with someone who genuinely cared about their work who wasn’t (or didn’t eventually become) competent. I actually think it’s healthy for an early-career engineer to sometimes get angry about their work, because it means they care a lot: it’s still a mistake in the moment, but it’s a “good mistake” . I certainly used to get angry — in fact, I wrote about the angriest I’ve ever been at work here 4 . But you have to move past it . Think of “caring about your work” as a vertical tube, unsealed at either end. You fill the tube by pumping in emotional investment from the bottom 5 . If you have too little, it drains away and you end up as a useless coaster. But if you have too much, it overflows and you end up as an angry engineer that people have to work around. One solution is to try and care the exact right amount: be invested in work a bit, but also have hobbies and a family and whatever else gives you perspective about your work problems. If you have a rich and healthy personal life, it’s hard to find yourself yelling at somebody about React state management. However, this is a tricky balance to maintain over time. Another solution is to care about different things. The reason too much caring overflows into anger is because what you care about is misaligned with what the organization cares about . If your interests are perfectly aligned with your company’s (for instance, if you primarily care about delivering shareholder value ), you can fit way more emotional investment into the tube before it overflows. Here’s some dangerous advice : showing a little bit of anger at work can sometimes be useful. It can be a good way to signal that you care, or to build rapport with certain people, or to draw attention to something you think is important. However, it’s still always a mistake to be angry. You need to be able to drop back to a friendly mode at will, which is very difficult when you’re genuinely angry. Being able to show a full range of emotion at work is good. It makes you more persuasive and more human. Being a fully professional robot is fine — you can have a successful career this way — but there’s always going to be some kind of uncanny-valley HR-ness to your work persona that will make it hard to connect with your colleagues. If in doubt, don’t show anger. It’s never wrong to be professional. However, if you can signal that you’ve got enough distance to separate your professional feelings from your real feelings, and enough perspective to realize that the stakes of a technical decision are fundamentally not that high in the grand scheme of things, it can sometimes be okay to show visible frustration so that people know you’re still human. Well-known software engineering personalities are often angry. It feels unfair to give too many negative examples, but obviously Linus Torvalds’ rants about Linux are a great example. Some of my favourite engineering talks are from Bryan Cantrill, who is sometimes visibly furious at his subject matter. There are too many well-known angry blog posts to list, but I’ll cite one I genuinely like: my Australian blogging colleague Nikhil’s post titled I Will Fucking Piledrive You If You Mention AI Again . Anger is a part of the general image of a competent software engineer. Many junior engineers learn from this that it’s okay to be angry. However, taking your emotional cues from engineering celebrities is a big mistake, for a few reasons. First, you are not Linus Torvalds or Bryan Cantrill . Torvalds is the BDFL of the most important software system in the world. Cantrill is the cofounder and CTO of his company. When these people are angry at work, people will not work around them, because they are the ones deciding what gets worked on . Once you’re the one in charge, you can get away with being emotional in the workplace 6 . Second, you don’t know what it’s like to work with these engineers . People give talks and write blog posts because they’re emotionally worked up about something. If your only exposure to a celebrity is via their conference talks and blog posts, you’re seeing them at something like their maximum emotional intensity. If you then take that level of emotion into your normal everyday work, you’re almost certainly overshooting. I’ve been reorged into dysfunctional teams, have had projects I enjoyed cancelled, and have worked on systems that were extremely chaotic. I can’t remember the last time I was actually angry at work. To be clear, I’m not successfully hiding my anger (unless it’s so repressed it’s invisible to me as well) 7 . Nor am I naturally a chill person. I’ve just reached a point in my career where I genuinely don’t get upset about work stuff. A cynical person might say here that I’ve stopped caring about my work, so of course I don’t get angry anymore. I’ve left the side of the “real engineers” — the Linus Torvalds and Bryan Cantrills of the world — and sold out for that sweet, sweet big tech money. I mean, maybe! It’s true that I’m less invested in specific technical decisions than I used to be. But I still care a lot about doing a good job, I still spend a lot of time tweaking and reading code, and I certainly get more done than I did when I was more emotionally volatile. Being angry at work feels good. It feels like proof that you’re working on something that matters, and that you’re personally having an impact. If you’re angry, nobody can call you a coaster. But anger is only a local maximum. If you can find your way to a different style of working, you’ll not only be more effective, but you’ll be in a far better place to have impact on problems that actually matter. Mostly, I fail . As I understand it, this is the work environment that most famously angry engineers came up in. My experience is certainly limited (a handful of companies, and maybe ten different teams or organizations). I can certainly believe it happens! About halfway down, in the section titled “it’s not your manager’s fault”. Emotional investment is a liquid with the viscosity of water. To a point. Even Torvalds famously said he’d gone too far with the anger and decided to turn it down a bit. I suppose I’m not the best person to judge whether this is true. If you work with me and I do come across as an angry guy, please do tell me. Mostly, I fail . ↩ As I understand it, this is the work environment that most famously angry engineers came up in. ↩ My experience is certainly limited (a handful of companies, and maybe ten different teams or organizations). I can certainly believe it happens! ↩ About halfway down, in the section titled “it’s not your manager’s fault”. ↩ Emotional investment is a liquid with the viscosity of water. ↩ To a point. Even Torvalds famously said he’d gone too far with the anger and decided to turn it down a bit. ↩ I suppose I’m not the best person to judge whether this is true. If you work with me and I do come across as an angry guy, please do tell me. ↩

0 views
Brain Baking 2 days ago

Deconstructing A 17th Century Letter

Continuing the letter writing fever, I’ve been exploring the academic work called Letters As Loot by Gijsbert Rutten and Marijke J. van der Wal. It’s a treasure trove of information on how letters were written in the 17th and 18th century, based on a corpus of 50k letters that were captured by raiding ships between England and The Netherlands. These boats usually contained a bag full of letters headed to offshore family and friends but due to the raging Anglo-Dutch War (four of them between 1652 and 1784), many letters never made it to their destination. That’s frustrating for the reckless explorers waiting to hear something from their family back home (or the other way around). But at the same time, nearly four hundred years later, these lost letters are the only thing that help us shape our view of lower and middle-class Dutch letter writers, which is otherwise dominated by upper class male writing. I’d love to think that if the Belgian Post loses a letter I sent, it’ll turn up in a couple hundred years for researchers to marvel at. I’m pretty sure by then this website is long gone. Who says the internet will still be alive? Or the planet, for that matter, judging by the extreme weather conditions… Unfortunately, Letters As Loot is written in a stereotypical form of Academese and hard to recommend in its current state. So I scanned the pages diagonally and summarised the coolest findings here. The most interesting is the deconstruction of a typical 17th century letter, written by Angenieten Cornelis to her husband at sea Roelandt IJosten Ostvorndijck. Here’s what she originally wrote: A letter by Angenieten Cornelis, dated 15 September 1664. Lifted from Letters As Loot, page 80. That’s quite challenging to read even for a Dutch speaking person like me. Here’s the line-by-line translation by the authors: What’s so interesting about this letter? The fact that more than 50% of its contents is what we’d nowadays call formulaic nonsense: If you’re married to Roelandt (written as Roellant, which I’m quite sure is misspelled), why write such a cold an distant contents? She’d “regret to hear it” if he wasn’t in good health anymore. I think my wife would kill me if I’d write like that, thinking there must have been a very suspicious reason for my aloofness. She must have made up for that by ending it lovely: a hundred thousand times good night by me —but just to be sure, write your full name again. It is I, Leclerc! Some expressions, such as a way of ending the letter, were so common that they were abbreviated: uwe dienstwillige vrient (your faithful/obedient friend) would become U.E.D.W. . The most funny one most certainly is ick sal onse kijndt eens soenen voor u (I’ll kiss our child for you). It only appeared a couple of times so the writer must have been a creative one. The use of a very formulaic structure was extremely common in the 17th century, which gradually disappeared in the 18th century and onwards. The authors of the research also discovered that this greatly varies between classes: lower class and female writers are much more likely to write like this, while upper class and male writers seem to skip this entirely. It hadn’t occurred to me that this is largely due to (il)literacy. Letter writers having difficulty writing will stick to copying example letters from letter-writing guides. The authors make the distinction between more elite manuals, the very popular school books, and Jacobi’s Gemeyne zeyndt-brieven from 1645 that was reprinted well into the 19th century. Most of these books came with example letters and offered formulas for starting a letter, expressing regret, how to end it, how to note the date, and so forth. Kloek en gezond (strong and healthy), laat u weten dat (to let you know that), blijft hiermede den here bevolen (stay with this the Lord recommended) and so forth are all examples from Jacobi’s book 2 . If we look back at Angenieten Cornelis’ letter, you can easily discern some basic structure and delete all the junk to filter out what Angenieten actually wanted to tell her very respected and devoted husband. Here’s how I imagine that letter might be written nowadays: How much formulaic sentence and paragraph structures do we use nowadays? Barely any, and yet, the main form of the letter hasn’t changed since the 17th century: there’s still an opening (date, addressing to), a short formulaic exchange of pleasantries ( hope you are well, how’s things ), the bulk contents of the reason for sending the letter (here’s some stuff), and a closure ( cheerio or as M.I.A. would write, XXXO ). As the time went on, the strict use of formulaic structure disappeared. Thank God for not having to thank God a thousand times good night anymore. Most letter writers, including me, would start a letter with the date & location, but most older letters would switch this up and end with it. Also, the address we now place on the envelope would be the beginning of the letter, integrated as part of the text. This, together with the formulaic weirdness, would make it easy to spot if a piece of paper with scribblings on it was in fact, a letter, and not some customs clearance notice, tax slip, or worse, declaration of Independence. In sum, for the Dutch history nut into letter writing, this bottom-up research of the “ Sailing Letters ” by Rutten and van der Wal is a gold mine. I didn’t even touch the importance of these letters to the etymological history of written versus spoken words in Dutch. The digital version is freely available. Don’t let the dry text discourage you; just scroll through it to find a cool photo and read a few bits here and there. Time to buy a good brown ink. Looft God Booven Al! Did you know that the address field didn’t need a number? “To cobbler Jeff, the one living near the town square” would suffice. Or “To my husband Jeff, at ship The Jolly Jumper at sea left via Amsterdam commandeered by captain DoGood”.  ↩︎ Later authors attempted to modernise Jacobi’s increasingly outdated manual, but as Rutten and van der Wal note, the “differences outweigh the similarities”: most typical formulaic sentences were copied from other letters, not from examples out of manuals.  ↩︎ Related topics: / letter writing / By Wouter Groeneveld on 21 August 2026.  Reply via email . God is praised in the “address field” 1 ? Why write worthy and very beloved husband to address and start the letter? Why repeat the full name of your husband and yourself? Surely your husband will know your name? Nothing more on this occasion than commend to the Lord? Did you know that the address field didn’t need a number? “To cobbler Jeff, the one living near the town square” would suffice. Or “To my husband Jeff, at ship The Jolly Jumper at sea left via Amsterdam commandeered by captain DoGood”.  ↩︎ Later authors attempted to modernise Jacobi’s increasingly outdated manual, but as Rutten and van der Wal note, the “differences outweigh the similarities”: most typical formulaic sentences were copied from other letters, not from examples out of manuals.  ↩︎

0 views
Stratechery 2 days ago

2026.34: App Snore

Welcome back to This Week in Stratechery! As a reminder, each week, every Friday, we’re sending out this overview of content in the Stratechery bundle; highlighted links are free for everyone . Additionally, you have complete control over what we send to you. If you don’t want to receive This Week in Stratechery emails (there is no podcast), please uncheck the box in your delivery settings . On that note, here were a few of our favorites this week. This week’s Sharp Tech video is on the turnover at DeepMind. Apple Makes Compromises in the EU. Ben has covered the angst surrounding the App Store since the beginning of Stratechery and was focused on Apple’s policies long before it was cool. Now that the company’s finally been forced to compromise in various forums — including a settlement this week with the EU, as well adjustments to its ATT policies in Germany — I thought the most remarkable aspect of Ben’s coverage on Wednesday was how incidental and boring it all seems in the shadow of the possibilities and concerns that exist everywhere else in tech right now. We had a fun conversation about that dynamic at the top of this week’s episode of Sharp Tech before turning to AI cybersecurity, vibe coding epiphanies, and more insight on writing with and without AI. — Andrew Sharp Truth (Social) and Reconciliation. Sharp China returned from its annual August hiatus this week, and in an episode that’s outside the paywall , we talked about various sources of U.S.-China friction before Xi’s visit to D.C. in September. Before that, however, we began in Korea with more questions than answers as Foreign Minister Wang Yi descended on Seoul in the wake of President Trump’s abrupt Sunday evening decision to reduce joint military exercises between the US and ROK. As for that Trump decision, in this week’s Sharp Text article , I used the Korea news as an opportunity to marvel at the exhausting economy of takes and theories that accompanies every foreign policy decision (and meme) under the current administration. — AS August Fun with the Clippers and Lakers . During the quietest period of the NBA calendar, there’s actually been quite a bit of news out of L.A. On one hand, we have a terrific mess as Buss family members squabble and Mark Walter’s DOJ-flavored cashflow problems have led to a shocking sale nine months after he initially purchased the team. On the other, Steve Ballmer and the crosstown Clippers might be in the (relative) clear after a 12-month NBA investigation into alleged salary cap circumvention. We discussed all of it on this week’s Greatest of All Talk , including frustrations with Clippers media coverage, what the NBA wants for the Lakers, and a memorable Top 5 segment about our top vacations.  — AS Stripe Acquiring OpenRouter, Aggregating AI?, Flipping the Business Model — Stripe is reportedly acquiring OpenRouter, an implicit bet on a future market of models and the chance at Aggregation. Nvidia Backs OpenAI Data Center, Anthropic News, Google Buys Spirit Airlines Data — Nvidia makes another deal, this time with a frontier lab; Anthropic’s revenue continues to amaze; and maybe data finally is oil. Apple Settles With E.U., U.S. App Store Fees, ATT Rules in Germany — Apple’s App Store is finally facing the reality of lower fees, and the EU should be satisfied with its work; it’s ok it’s late. So What Was Trump Saying to South Korea on Sunday? — A snapshot of Truth Social foreign policy and the take economy it inspires. More on Watermarking Apple Settles With EU How TSMC Uses Old Fabs to Make New Chips China Built 700 Waste-to-Energy Plants in 6 Years Wang Yi Visits South Korea; Remembering Zhu Rongji; US-China Ahead of Xi’s Visit; How China Monitors Foreigners August Fun with the Lakers and Clippers, Top 5 Takeable Teams or Players, Top 5 Vacations The App Store in the Shadow of AI, Offensive and Defensive Cybersecurity, Q&A on Financial Planning, AI Writing, American Sports

0 views
Manuel Moreale 2 days ago

On values, morals, and doing business

I caught wind of the news that MacStories is back posting on X, the “everything platform”, and that seems to have caused some backlash. I am not a reader of MacStories (because tools are tools and I don’t read about new computers the same way I don’t read about new washing machines) and I’m also not on X. So why do I even care about this news? Well, the short answer is that I don’t. But after the move (I guess because of the backlash?) Federico Viticci, MacStories’ Founder, posted what he called “ An Explanation ” conveniently, not on MacStories’ homepage and not even on the site at all. You should go read the whole thing if you’re curious, but there are a couple of things here and there that caught my attention. The first interesting thing is the attempt to separate Elon and his morals from X. We want to be absolutely clear about something: returning to X does not represent a change in our values, what MacStories or we stand for, or an endorsement of Elon Musk. We abhor Musk’s politics, his rhetoric, his careless disregard for others, and the direction he has taken X. But we also disagree that participating on X means we have aligned ourselves with any of Musk’s views or actions. We reject Musk’s worldview and will continue to treat people with the kindness and respect they deserve wherever they are on the Internet. I’m sorry, but this is not how things work. By being on the platform, by using it to reach the audience that’s on there, you are directly supporting the man. Because the platform only has value because people are on it, and you being there makes the situation worse, not better. The main reason why MacStories is back on X is, surprisingly, money. Boring, yeah I know. Quoting from the post again: It’s not just how we earn a living; it’s one of a small number of independent websites that still cover apps, Apple, and a growing list of topics, including videogame hardware and the automation and productivity side of AI. Readers shouldn’t have to think or care about the business side of MacStories, but we have to, which is why we returned to X. If that is the situation, be transparent. Show how you run the site, how much money the site is making or losing, and make your case. But simply saying “Sorry, running a site is hard so we’re doing a 180 without telling anyone, not even the people who work here” is a shitty move. Especially because the move is pretty obvious: the amount of AI-related content on MacStories has exploded, AI people are on X, and if you want to try to grab some of that money, you have to play that game. Like, I get it, but at least say it and own it. Don’t try to play the running-a-business-is-hard card. Have the guts to put up a post on your precious website where you explain what you’re doing and why you’re doing it. Tell people you want to chase more AI content, because you think you’re now more of a “builder” . You’re an Apple fanboy; have some courage. People with strong values and morals seem to become rarer and rarer these days, especially when money is involved, and it’s so sad to see. And this whole AI moment seems to be exacerbating that. Thank you for keeping RSS alive. You're awesome. Connect via email :: Sign my guestbook :: Support for 1$/month

0 views
Unsung 2 days ago

Movie review: General Magic

★★☆☆☆ 2018, 92 minutes General Magic was a company started in 1990 by some of ex-Apple staffers, working on a pocket communication device and its operating system (Magic Cap, sporting a fascinating room user interface). The product launched in 1994, but swiftly failed in the market, similarly to and concurrently with its chief competitor, Apple Newton . Many alumni of the company – including Megan Smith, Kevin Lynch, Tony Fadell, and Pierre Omidyar – ended up having successful careers in tech after that, to a point that General Magic is referred to as “ Fairchild Semiconductor of the 1990s .” = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/movie-review-general-magic/1.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/movie-review-general-magic/1.1600w.avif" type="image/avif"> The eponymous documentary , released in 2018, was disappointing to me. Sure, it is well-edited, and particularly shines owing to a lot of archival footage shot during General Magic’s happy years. Unfortunately, we get to see none of the sort of details I was hoping to see: no pre-release interface or hardware, no discussion of product nuances, no solid reflection on the company, the culture, or the zeitgeist. The movie drops so many fascinating questions and avenues, but doesn’t really follow up on them: Unfortunately, I kept thinking of the movie as “generic magic” – a sort of interchangeable valorization of Silicon Valley’s “failure is secretly success in the fullness of time,” defaulting to romantic or rousing music, that must have felt obsolete even in 2018. It felt like the documentary could very well talk about one of the many other companies and efforts, which is frustrating, because there was something special and unique about General Magic. I also kept remembering The Soul of the New Machine , Andy Hertzfeld’s own Folklore.org , and even Halt and Catch Fire , all of which more adeptly interspersed personal and emotional drama with specific details you could learn from and take home with you. I think General Magic deserved more. (I do want to recognize that perhaps this wasn’t a movie for me. The documentary is free on YouTube if you want to check it out – it’s consistent throughout, so if you like the early minutes you might like the whole thing.) = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/movie-review-general-magic/2.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/movie-review-general-magic/2.1600w.avif" type="image/avif"> (To play with Magic Cap, go to Infinite Mac , click on Macintosh Garden tab at the bottom, search for “magic cap”, click on magic_cap_simulator.sit, go back to the emulator, double click on the Outside World icon, double click on Downloads, double click on the .sit file, wait for it to unpack, close the Downloads window, reopen it, double click on the Magic Cap Simulator icon.) #apple #history #movie review #review At some point, Andy Hertzfeld reflects on setting a bad example by focusing on small playful details instead of rallying the team to ship stuff. Where did this realization come from and did that change him as a person going forward? In hindsight, did the company benefit or suffer from the culture of freewheeling “superstars” who are also perfectionists? In a strange, brief vignette, a few people talk about not wanting to have any managers – but that is not picked up again or resolved in any way. How did the company manage to hire all of this past and future talent? Are there any repeatable lessons in here? During the launch, there is a brief slide with “whole person thinking,” which felt unique for a tech product reveal. Near the end of the movie, Kara Swisher talks about worrying how mobile devices can affect and perhaps even impair human-to-human communication. Those two threads are not connected. About the only lesson learned and spoken out loud comes from Tony Fadell, who says: “with the iPod, we iterated a lot faster.” Would this have mattered if the premise of the movie seems to be that General Magic was too early on the market anyway? There was a mention of a skeuomorphic room interface done pretty much overnight by Andy Hertzfeld. From the perspective of today, it is the device’s perhaps most distinctive characteristic, but it’s not covered more than that one mention. In another very rare specific example, there is a beat talking how iPhone’s (very controversial, early on) software keyboard owes its existence to General Magic’s team trying that first. How did whoever created it feel about it? An external observer suggests that General Magic missed the early ascendance of the web, but we don’t see anyone from the company reflecting on it.

0 views
matduggan.com 2 days ago

In Purgatory Everyone Likes Crepes: My Time at a Greek All-Inclusive

I can’t remember the last time I was actually hot. It’s a thought that keeps looping through my head as I stand next to a children’s play structure in a Greek all-inclusive waterpark. The sun is beating down on an international community of mostly parents drinking watered-down cocktails out of plastic cups while their children scream and run in circles. Surrounding this little paddock of plastic is a lazy river, filled with mothers and fathers staring blankly into the middle distance as they slowly orbit the water toys. It’s all concrete and aggressively bright primary colors, suffused with the smell of gyros being cooked by a Filipino staff behind me. Living in Denmark, it’s rare for the sun to shine directly on you. Typically, everything outside looks like it has a blue filter placed over the lens—like how TV shows put a yellow filter over the camera to establish that you are now in Mexico. Sometimes in Copenhagen you might get hot, but then it will immediately start to hail on you, and you suddenly have a different, worse problem. Greece, by contrast, feels like the sun is noticeably closer to the earth. It feels personal, like a heat lamp pressed against your skull. My daughter has been playing peek-a-boo with a cute baby on a bench next to this water playset for long enough that I feel a moral obligation to go over and say hi to the mom. She’s a nice woman, watching four kids at a waterpark alone with a casual calm I actively envy. I’m a little stressed keeping eyes on just one, but she has four kids ranging from around ten to two, running around and disappearing beneath the waves, only to emerge just before I would panic. She has a small sunburn on the small of her back that I assume she either can’t reach or forgot—a small, raw triangle of red. We chat briefly. She asks where I live; I ask the same. I’m surprised when she answers Russia, mostly because I don’t get much exposure to Russians in Denmark. I clearly make a face without meaning to, because she quickly adds, "Their father is off fighting in the war." Ironically, I think she added this to head off me judging her as a single mother. Instead, revealing the father is a Russian soldier has made the part of my brain that regulates politeness completely short-circuit. What is the social etiquette for the wife and children of a military power you don’t support? I’ve never even considered this question, standing here in my swim trunks, my too-pale stomach exposed to a hostile world. I don’t want her children to be orphans, obviously. Neither she nor they did anything to start this conflict. But I’m worried that expressing vague positivity will end like my conversation with an Israeli in Copenhagen who took my neutral statement, "I hope the conflict ends soon," and responded with a wink and, "It will, with Trump at the helm." "I hope he comes home soon," I say, keeping my face perfectly neutral. She nods and resumes throwing a ball around with her children. She seems completely unaffected by thinking about her husband, which I think is actually the strange magic of the all-inclusive resort. It is a geopolitical anesthesia. The world is burning down, but here, the ice cream machine is still running. On most topics, I don’t consider myself a snob. I enjoy a Michelin-star restaurant, and I enjoy a nice slice of 7-Eleven pizza at 2 AM. I'll finish a thick, sad Classic of Literature and then immediately pick up Space Marines vs. The Metal Skeleton Wizard on the Moon . But on travel, I’ve always looked down on the all-inclusive crowd. Sitting one rung above a Carnival Cruise on the ladder of human decay, the all-inclusive always seemed like the most cynical way to say you had legally stepped foot in another country. My first real exposure to this was in Puerto Rico. I would leave my bespoke 20-room boutique hotel and scoff at the barbed-wire fence surrounding the American mega-resort down the road. "Why would you fly to the Caribbean to eat at a steakhouse?" I’d say, dripping with scorn to my wife, who would nod as we walked through dark neighborhoods on our way to eat dinner in someone’s actual living room that Google Maps called a restaurant. As I traveled the world, I saw these giant compounds everywhere. I swore I’d never become one of those placid, sunburned people wandering around a sculpted grass lawn, slightly buzzed at 11 AM. So imagine my surprise when I found myself finishing my second tequila sunrise before lunch, staring at a too-blue pool. The resort’s "activity crew" was out on Crepe Island—so named because it had a full crepe station on an island in the middle of the pool. "Really, it's a crepe peninsula," I say to nobody. Nobody cares. They are aggressively not paying attention to anyone else. This is not a community. A line of unlicensed, Disney-and-Nintendo-adjacent costumes are paraded out, with "Merio" closest to me. Kids kept trying to give Merio a high-five. I don't know if it was the suit or if the guy inside was just sick of his job, but he kept waiting too long to high-five the kids back, inevitably wacking them in the face or the back of the head as they moved on. Over and over, an excited child would run up, followed by a bizarre, lagging delay before the inflated hand would strike. "At some point, that's just hitting kids," I said to the ether. The woman next to me put her AirPods back in. My wife had booked this trip as a reward to our daughter. Having been a remarkably good sport as we dragged her to such famously child-friendly vacations like "A Sheer Cliff Overlooking Icebergs in Greenland" and "Endless Climbing Up and Down Stairs in Florence," we wanted to give her a chance to just be a kid. When we asked her what her favorite vacation was, she responded with zero hesitation. "Lalandia," she’d say, before resuming coloring in her coloring books with so much force you could see the head of the marker get pushed back inside the plastic casing. Since Lalandia is effectively a massive, indoor all-inclusive in Denmark, we figured the same concept in Greece would at least be a warm alternative. We boarded a chartered flight direct to Greece, my first time on a flight where everything was delivered in Danish. I felt like a spy hiding among the locals, fearful they'd discover the American among them. Safe from prying eyes, I was pleased to watch the Danes engage in behaviors they usually made fun of us for. Loud conversations were had, crunchy snacks were aggressively unwrapped, and general uncouth behavior typically shunned was suddenly fair game at 30,000 feet. Upon landing, the tour coordinator told my wife and me that we had to "run" to make our bus. My wife took off for the toilet; I grabbed the kid and the big suitcase and sprinted for the van. We then sat in the van waiting to take off for over an hour. What I hadn't realized in my heroic rush to the van was that I had left a small carry-on sitting in the middle of the airport. It was filled with snacks, coloring books, and also my wife’s wallet and passport. The first night at our all-inclusive was understandably tense. Variations on the word "idiot" were thrown around. The next day, we got the good news that the bag was safe and sound at the airport police station on the island of Kos. So while my family went off to the pool, I went to the lobby and called a cab. My driver arrived 30 minutes later and immediately asked, "Is it ok if you sit up front so we also take these beautiful people downtown?" The Scottish woman in the back blushed at this. I, lacking the social architecture to recognize warmth, simply thought: Wow, bold move saying that in front of her husband. I got into the passenger seat. This, as I realized later, totally broke down the driver-passenger social construct. After dropping the charming Scottish couple off in the downtown area, my driver started telling me literally everything about his life. For the next forty minutes, I learned how his brother had been a firefighter, then one day went on vacation to Vietnam, fell in love with a Vietnamese woman, and opened the only authentic Greek gyro restaurant in Hanoi. He had a booming, theatrical laugh and was poured into a skin-tight polo shirt. "Why are you going to pick up this suitcase your wife lost? DO WE EVEN NEED TO ASK?? HAHAHA! WOMEN!" Then he’d get deadly serious and remind me, "Children, they're the most important thing." I had spoken maybe six words at this point in the trip. I forced a laugh, silently calculating the statistical probability of being murdered and buried in a Greek olive grove. It would be a beautiful place to die, though I worried they’d use a photo of me from this trip—pale, squinting in a baseball cap—for the Netflix documentary thumbnail. Some true crime podcast would inevitably conclude I had it coming. At one point, he pulled the cab over to show me a flat patch of grass. "This is where I saved the island from a wildfire," he said proudly. I couldn't see any sign of fire; it just looked like a normal field with the grass crushed. I didn't say anything, as up to this point, the conversation had been delightfully one-sided. I had mostly sat sweating through my shirt, nodding. He told me how someone had thrown a cigarette out of their car, and it had started to smolder. He had stopped and used his legally required fire extinguisher to put it out. "It's good that I had this," he said, patting the dash. "This is why a robot can never replace taxi drivers." He then stared at me as if I were an assassin sent from Silicon Valley. "You betcha," I said, sweating. I wanted to assure him I wasn't there to replace him with an app. My amoral peers back home were already working on it, but I personally wouldn't kill his livelihood. Just people who looked and sounded exactly like me. I kept poking around silently instead, but then he caught me looking at the giant bulldozer sitting twenty meters away, its tracks leading directly to the field. I'm a genius like that. "Yeah, alright, also the bulldozer helped a bit to put the fire out. But mostly, it was my fire extinguisher." The bulldozer was merely an accessory after the fact. He then changed the subject by telling me that Tom Hanks had sat in his taxi "right where you are sitting." "Did you know Tom Hanks loves Greece and we love him? He is such a friend of Greece that we gave him a passport." Is Tom Hanks Greek? I thought, staring out the window. Why are there so many interior decoration stores? Is everyone on this island redoing their bathroom? "Tom is so nice, we love him." My brain returned absolutely no information about Tom Hanks, which frankly is typical of my mind. He clearly wanted me to respond with a Tom Hanks factoid, but all I could think to say was, "Yeah, his son was really good in Fargo ." "HIS SON? Who is talking about his son?" He took the conversation back over, explaining how Tom Hanks’s wife had sat in the back of the cab while Tom sat in the front. Weird choice, Tom, I thought, looking for a sign we were wrapping this trip up. When we arrived at the small airport, he parked randomly outside the main building, basically just the corner of the road that led from the airport back to town. "Wait, are you coming in with me?" I asked, confused. He nodded, saying it would "take way too long" if I went in by myself. Suddenly we were concerned about efficiency. He marched us in, knowing everyone who worked there. I found myself in a small police office that definitely didn't understand the separation of church and state. There was a massive crucifix on the wall, and each police officer's desk was covered with a Jesus mousepad and a large picture of the Virgin Mary. My cab driver and the police officer negotiated the release of the suitcase in front of me without bothering to involve me. At some point, I was told I could take the suitcase and get out of here. Nobody had asked me any questions except to quiz me about the contents of the bag. On the way back, we stopped for tea because, truly, what is a taxi meter at this point? We'd been in the car together for like two hours. He told me how his mother-in-law had almost died from a heart attack. "It's good that she could get to Athens in time," he said. He then told me how during the off-season he harvests olives from his family farm and has them pressed down the street. "It's heaven working the fields with your family." I imagined my siblings and I being asked to harvest olives in the beating sun and immediately envisioned four body bags in the shade of an olive tree. I nodded, staring out the window at a fire extinguisher store across the street, where what looked like a 12-year-old boy rolled a cigarette seemingly one-handed and lit it while sitting on top of a barrel. The long grass around the fire extinguisher store seemed primed for a fire. What happens if a wildfire hits a fire extinguisher store? I pondered, while my driver went on about the beauty of the Mediterranean or something I wasn't paying attention. Then we packed it up and went back to the hotel. "Thanks for a nice morning, Mark," he said fondly to me as I got out. My name is Mat. But I guess if you've saved the island from a wildfire, you can call me whatever you want. I got back just after Mini Disco. This was a one-hour dancing marathon for the kids in the semi-enclosed theater. With two fully manned drink stations on either side of the stage, it was an opportunity for children to dance while their parents drank like an asteroid was imminent. The songs stayed the same every night, but on this first night, we were going to learn a valuable lesson: Always buy the t-shirt. Apparently, while I was learning the ins and outs of the Greek cab business, a nice woman had asked my wife and daughter if they wanted to buy a t-shirt for the Mini Disco. My wife, smelling a tourist scam, had hard-declined. What she didn't realize was that the climax of the Mini Disco was a formal t-shirt presentation ceremony. Each child's name was called, they were presented with a t-shirt, and then they were allowed to hug Leo the Lion. My daughter didn't understand that she didn't have a t-shirt. So, when the name "Eleanor"—which is not her name—was called, my daughter hopped up and snagged the shirt. This meant my wife had to rush onto the stage, rip the t-shirt out of our daughter's hands, and hand it to the sheepish little girl whose shirt it actually was, all while holding a cocktail. It was a masterclass in parenting under the influence. The next day, we were on the hunt to put in an order for the t-shirt, having made the most serious promises parents can make to a child. After we chased down the red-haired French woman running the kids' merch table, we went to the main pool. She didn't seem surprised to see us come crawling back. I got the sense the first day she was doing the sales pitch, then after that she just waited for us to come to her. There is a weird etiquette to pools and British people. They will rush out the second you are legally allowed to put a towel down on a chair, meaning from 8 AM to 1 PM, there isn't a single open seat. The British treat pool chairs the way trench soldiers treated no-man's-land: as hotly contested territory worth dying over. But if you are lazy like my family is, you just wait. They start to leave the pool around 1 PM, and you can get a great seat with zero work. The next day, you repeat the entire cycle. That evening, kiddo got her shirt. It was a proud moment; she clutched her blaze-pink t-shirt to her chest and teared up with pride. Soon, the days started to blend together. At some point, we all got an emergency text message telling us about wildfires on a neighboring island. I assumed the resort staff would need to calm down the crowds, as you could pretty clearly see the smoke from the fire. The sky was turning a dark, orange-brown in the distance, and the air smelled like burning rubber. Nobody cared at all. A thousand people looked at a Greek text message they couldn't read, put their phones down, and picked up their beach reads. The apocalypse was met with a shrug. As an allegory for global warming, there is something particularly bleak about people dancing in a pool to 90s boy band hits as a thick plume of wildfire smoke shoots up into the sky. Thankfully, here nobody pretends to care about anything, so I was able to slip back into a soothing apathy. In a world that was constantly asking me to pay attention to some fresh horror, this was a place that asked nothing of you. Toward the end of the trip, we decided to head down to the small town and were dropped off at Dolphin Square. We started walking around, looking at Google Maps to see what we should look at. There was a Roman House, which is basically an empty lot full of pieces of old Roman architecture and a sign saying I needed to pay 10 euros a person. Since nobody was there to collect this 10 euros, the sign felt more aspirational than practical. Someone should pay this Greek island 10 euros, but not today, I guess. But I looked at the old rocks. My daughter was unimpressed. She looked at the ruins, looked at me, and asked if she could get ice cream. "Denmark has a lot of old stuff," she said, which, in her defense, is true. We walked around a bit more, buying little tourist trinkets. I'll never understand who buys the t-shirts that fill the small streets of these towns. Who is the demographic for a shirt that says "I heart my boyfriend"? And more importantly, where is he? Does he buy the t-shirt for you before you go on a solo trip without him? But we went to a small restaurant and had a proper meal for the first time in days. Despite the heroic efforts of the staff at the resort, it was actually hard to eat there meal after meal. The culinary philosophy seemed to be "cook the will to live out of it." Everything was either fried and dried out beyond belief, a hard puff pastry, or so bland you couldn't really tell what was going on. The lamb sorta tasted like the chicken, which tasted a lot like the pork, which tasted like a cry for help. At one point, we went to an Asian-themed restaurant on the resort, and I was served "duck with Chinese pancakes." The duck was well-cooked, but the pancake was an Old El Paso flour tortilla. I couldn't help but feel like the duck died for no reason. At one of these meals, where my daughter pounded another plate of french fries and pizza, a British woman at the table next to us told me how her entire family looked forward to this trip. She started to tear up as she explained that she and her husband had worked extra hours this year to make it so their kids could come with them. "It just makes me burst with pride to think about," she said. As she spoke, one of her kids sat beside her, wearing noise-cancelling headphones, staring at an iPad, and methodically eating an entire pepperoni pizza without making eye contact with another human soul. By the end of our week, I was eating meals of plain bread with watermelon and coffee. My daughter was consuming what looked like a kilo of Nutella, and my wife and I settled down for another cycle. We only ended up going to the actual beach, which was beautiful and maybe 200 meters from our hotel room, once. The ocean was beautiful, with rolling hills in the background, perfect warm water that was the saltiest water I've ever been in. It was like swimming in a giant, warm tear. Everyone but me hated it because it wasn't as nice as the pools. "Ugh, there are ROCKS!" my daughter shouted, upset at the audacity of the ocean for existing. Honestly, the rocks hurt a fucking ton, but I was too self-righteous to admit it. "Let's enjoy nature, everybody!" I yelled, bleeding from the feet. I thought by the end of my time with these people that I would come to think less of them. I expected to leave despising their complacency. Instead, I found myself weirdly protective of these folks. They are just trying to make some childhood memories with their kids, trying to manufacture something, anything that looks like a normal childhood in a world burning down. I won't lie, the all-inclusive life isn't the life for me. It is a place out of time, where the enjoyment of it requires a total suspension of your relationship to the outside world. But I do now understand the appeal of not being challenged. In a time when everything is being questioned and every tradition and norm is falling apart, this is a callback to frankly an easier time to be alive. And we ended up getting the t-shirt, which, when you break it down, is really the most important part.

0 views

I vibe-coded a C compiler that can build SQLite

A while back, Claude Opus built a pretty ambitious compiler. I didn't set out to do anything nearly as ambitious as Anthropic's. They were targeting multiple CPU architectures and they wanted to be able to compile a bootable Linux kernel. (I'd actually forgotten about that project until I started posting some screenshots of the work below to social media and someone reminded me.) Last night, as I was getting ready for bed I was reaching for some project to support and tool use in Evener, our ~new agentic harness. I popped open the mobile UI on my phone and typed "Your job is to implement a standards-compliant ARM64 C compiler for macOS in Swift." Evener asked me a couple of questions. I clarified my intent a little bit: "I'm great with radical task decomposition. You should use recursive subagents to manage context and complexity.Should structure the project in whatever sane way you want. You should work fully autonomously And do not need to ask me questions." And then I set a goal: "Implement a standards-compliant C compiler in modern Swift. It should be able to build SQLIte and have SQLite pass all tests. You may decompose the project in any way you see fit. You should use subagents, including recursive subagents, to execute effectively." I don't know what I thought was going to happen. I watched long enough to make sure that it was going to start working. At about four in the morning I rolled over and looked at my phone and it was deep into debugging a malloc issue. I woke up at a little bit before seven and it was implementing variadic functions. Today I spent most of my day in meetings and so I didn't get a lot of time to watch it, but I would occasionally pull up a session and see it making little bits of progress. Occasionally it would dump out some assembler and then dump out 's version of the assembler to hunt for differences. I saw it trying to build the 274k amalgamation and segfaulting. I saw it writing little test scripts. I saw it get to the point of trying to do an and exploding. When I came up for air at about 8 30 p.m. I popped open Evener around my phone, and was incredibly disappointed. My loop had just...stopped. It's not supposed to do that. But that's exactly the failure that I was testing for. I started reading the recent parts of the log, to figure out which obvious problem had caused the failure. I was not expecting to see this: It stopped because it had compiled SQLite and was able to do an and a . It did exactly what I asked. It took Evener + GLM 5.2 about 21 hours to build a C compiler capable of compiling SQLite and passing a basic smoke test. This is the checkpoint: https://github.com/obra/toy-c-compiler/commit/456ddfa8912b5ba47b968773ecddcdca741f0df2 (I was initially worried that it had cheated by doing the obvious web searches, but discovered late in the day that I had accidentally broken the harness's web fetch tool. And so it wasn't able to cheat that way.) Checking now, it looks like it didn't even try , which makes me happy. There are a ton of missing features. It is not yet standards compliant, but I didn't properly set the goal to force that. I've just kicked off another to get it to run through compliance suites and fill in the missing features. I'll probably let it run for another day or two just to see what happens.

0 views
matklad 2 days ago

Rust Glancer

Rust Glancer , a functional LSP server for Rust which uses two orders of magnitude less RAM, is incredibly cool. Go check it out! This post started as a comment on lobste.rs, but I figured it out that it’s better to publish it somewhat more prominently. Don’t expect polished writing though! Some thoughts: rust-analyzer uses rowan for syntax tree representation Yeah, rowan is garbage :P I was really thinking about And Rowan is pretty good for that. But that’s 1% use case. The 99% use case is all the code in your 6666 dependencies which you won’t ever look at, but which needs to be at least shallowly analyzed. Even for incremental tool whose main goal is refactoring, the primary AST structure should be just a list of arrays. There might be a real post about that at some point, see https://youtu.be/G93oYL1ry70 as a teaser. Rust workspaces genuinely have a lot of information that must be indexed: thousands of functions, structures, traits, relationships between these, function bodies and statements in them, etc. Each of these needs to be analyzed and remembered, and you can’t really cheat if you want to have things like “find all references to this structure”. If I understand correctly, Rust Glancer wants to process each function body. I think that part can perhaps be made lazy (but not incremental!) with little overhead? Index all items, but, for functions, do only the currently opened file? This might combine some of the better parts of both worlds. Would be interesting to compare memory usage with Rust Rover. Net of the IDE GUI itself, I would expect RR to be more compact. Some features are unlikely to be supported though, such as build scripts / proc macros support via proc macro invocation I might be rationalizing/misremembering things, but IIRC it’s exactly around adding proc macros that the thing began to feel unreasonably bulky. Expanding proc macros is slow as we are running real code, we can’t really do normal IDE cheats. And proc macros generate a lot of code. At one point I measured, it was like 30% of rust-analyzer binary size was attributed to JSON parsing code. If no one sees the code, it can’t harm anybody, right? One potential approach here is to pull the Sorbet trick, where you don’t run meta programming at all, and instead have a plugin interface to “explain” the effects of what that would have done. Instead of running serde, we just add a shim that injects with an empty body. I’m not sure why, but in rust-analyzer I’ve observed that when agents edit the code, inlay hints can get out of place Rust analyzer’s core data model is very pedantic about always observing consistent snapshots of the code, and does its best to ensure that the language client and server have a shared, strictly serializable view of the world. It’s a shame that LSP doesn’t allow that to be correct , only heuristically right , unlike the older Dart Analyzer protocol, which has sound data synchronization. However our implementation of file watching is sketchy! First, there are two backends: we can ask the editor to do watching for us, or we can use server side watching. Try changing this option and see if it helps? But then, yeah, my recollection is that our native watcher’s API was fundamentally racy, and I didn’t do the messy platform-specific work of making it correct. But the main thing I want to write, and why I moved from the cozy lobste.rs text area to the luxurious comforts of an Emacs buffer, is that right now rust-analyzer is a bit like that half-drawn horse meme, except that it’s only the head half of the horse. One Big Idea of IntelliJ is that it’s PSI API (essentially AST with resolved types) is really an interface, and there are multiple provides. And in a typical usage, there’s at least three backends in play: This is how I think such things should work. rust analyzer shouldn’t use salsa for all those 6666 dependencies you still haven’t looked at. It should just use rustc’s .rmeta files, switching to salsa, transparently, only when the user starts messing around their folder. The prerequisite for that is defining the abstract API for accessing Rust code. That was always the plan, and we did start on that at some point: https://hackmd.io/ytd82QNiT_Ku2XFr1EAtiQ rmeta-transparent – source code might not be available for some crates, the API should support pre-compiled rmeta files as inputs. But I don’t think that work was ever completed. This still seems to me to be the lowest-hanging watermelon here — split the world into arcy-pointy incremental tip of the iceberg, and mostly read-only, on disk, compact, dark, moist breeding ground for supply chain attacks. Such glance analyzer architecture would be great, imo! incremental parsing, incremental, DOM-mutation style refactorings, For the files opened in the editor, actively modified by the user, the PSI is backed by the concrete syntax trees. For the rest of the project files, the PSI is backed by the so called Stub Tree, a compact on disk representation storing only the “externally visible” parts of the file (so, without function bodies). If the user navigates to a new file, its PSI transparently switches from stubs to syntax tree. For dependencies, the PSI is often backed by the compiled .class files, produced by javac. If you navigate there, the IDE just decompiles stuff four you! Super cool!

0 views
Justin Duke 2 days ago

The Man Who Knew Too Much

The Man Who Knew Too Much is a film of peaks and valleys. There is little argument that it is a good film; whether you place it amongst the highest of Hitchcock's work depends on just how much you value the stratospheric heights it reaches, and whether they overcompensate for the lulls. The film is never weak, but it is weak for Hitchcock in parts. Watching it for the first time, and knowing that it is a Hitchcock film, you feel as though you don't quite need all the foreshadowing and lampshading. As it progresses, you can't help but feel a sense of déjà vu with his other work — remarkable for this film as compared to the parts of his filmography not considered his absolute best. As a remake of his own 1934 work, and something relatively mid-career for him, it is very much him in his bag, doing what he does and what he is well known for. Compare this with, say, Family Plot , where your mileage largely depends on how willing you are to go along with him doing something much shaggier than usual. Put another way: this is the first Hitchcock film where I found myself not once, but twice , checking how much time was left in the movie and being surprised that there was so much still to go. But those dizzying heights. There is, crucially, the ability Hitchcock has to wring noteworthy performances from the Jimmy Stewarts of the world — to see something new and interesting in a face you have already seen so many times. I say this without consulting any of my prior notes, but my gut reaction was that this was my favorite Jimmy Stewart performance I have ever seen. He plays a man who fits perfectly the definition of fumbling . Hitchcock uses his frame for great gags, especially in the first act. But there are two scenes in particular where you see Jimmy Stewart as something an actor of his caliber is almost never shown as: a completely broken man. First, the close-up of him holding the corpse of a dying man, trying to process on many levels the mistakes and errors he has made over the past few days. And then a scene that is grotesque in many ways, but not unrealistic, when he coerces Doris Day — his wife, who must be said is consistently much smarter than him throughout the film — to take a fistful of tranquilizers before he delivers the news that their son has been kidnapped. For all the love I have for Cary Grant, he tends to have a certain problem in Hitchcock films of being too starkly charming in the scenes where terrible things have happened to him. Jimmy Stewart crumbles, and stays crumbled, until the very end. And then, of course, there is the VistaVision and the scene work. I am not smart or well versed enough to talk about the technology and how Hitchcock used it, but the aesthetics are perfect and gorgeous. He relies a lot on static shots, especially when we are first introduced to Marrakesh, that — like a discordant note held for slightly too long — usefully upset the viewer. But where I must end this review is, I imagine, with the set piece that most people think about when they think about this film: Albert Hall. Five minutes of opera and those aforementioned static shots, not a single line of dialogue, but stress building and building and building — for me, it is indelible, and so powerful that even after I forget the silly international spy bits, I know I will remember it for years to come. I think this is correctly placed in the mid-tier of Hitchcock's canon. His best films are tauter and have more things to say, tauter and carry more opinions than this one does. But a middle-tier Hitchcock film is still a very good film. And if I got to watch something this good every evening for years to come, I would consider myself a very lucky filmgoer. 8 out of 10. One last thing: Doris Day was great. The script did not really give her much to display any sort of range, despite the character herself being very competent and capable and, again, much smarter than Jimmy Stewart's — but there's just not a lot for her to do. 1 Honestly, the same could be said of Jimmy Stewart's character, and to a certain extent that's why parts of the film drag as much as they do: you get the sense that our protagonists are, up until that Albert Hall scene, passive observers more than they are agents of the action. Not a sin in and of itself, but it exacerbates the drag in the middle of the film, because everyone is just going through the motions. But Day's performance in that tranquilizer scene is unimpeachable — and, as with Jimmy Stewart, not something you're used to seeing from an actress of her stature.

0 views
Sean Goedecke 2 days ago

Readers can't identify watermarked AI text

In the last few weeks, I’ve been complaining that everyone is wrong about AI watermarking: it isn’t really anti-consumer and it doesn’t make the outputs any worse. The watermarking papers demonstrate 1 that this is true, but I thought it might be interesting to put it to a practical test. Given examples of watermarked and unwatermarked answers to the same prompt, could readers tell which is which? To find out, I vibed up 2 https://sgoedecke.github.io/watermark-quiz/ , a static site that quizzes readers. I used Qwen3-30B-A3B-Instruct-2507 on a rented H200 to generate thirty responses: three responses per question, one of which was secretly watermarked with SynthID-Text. The rented GPU cost around two dollars. To measure results, I just sent users to a different page for each score, and aggregated visitors-per-page in my analytics 3 . This would be easily spoofable if anyone cared enough to do so, but for a casual test I think it’s acceptable. The first round of traffic I got to the quiz (278 participants) had these slightly puzzling results: Pure random choice would lead to an average score of 3.33/10. However, the mean score here is 3.92. There is indeed a spike around 3/10, as expected, but there’s also a second weird spike at 6/10. Why is that? It turned out that the SynthID response was option A in six of the ten questions, so users who just selected the first answer for every question would get 6/10. Oops. I re-shuffled the questions and got these results: Now the mean is 3.4/10, much closer to the expected 3.333. There’s no spike around 6. We only had 73 people take the quiz after I shuffled the questions — most people saw it and took it immediately after I posted it to my LinkedIn and Hacker News — but given the previous results, I think that’s still enough to feel confident that people were just guessing randomly. So no, people can’t identify the presence of AI watermarks . Obviously this wasn’t exactly a scientific study, but it’s still pretty suggestive. If watermarks were really choosing random words that the model would never pick, you’d be able to sometimes tell from three side-by-side responses which one went down the weird watermarked road, right? I also hope that something like this can serve as a persuasive tool: if you’re worrying about what impact watermarking is going to have, and your intuition is unmoved by the mathematical explanations, having a read of the watermarked and unwatermarked responses might convince you that there’s really no difference in quality. The one-sentence explanation for why is that AI models already randomly select from a handful of top tokens, and watermarking just replaces that random choice with a bias that is predictable while still being equivalently “random”: as a simple example, instead of “pick randomly from the top three tokens”, you could do “count the letters in the previous ten tokens, take mod three, then pick that token”. Some notes from the vibing: GPT-5.6-Sol put extraneous text all over the page I had to get it to remove, it chose the now-very-recognizable styling that I had to rip out, and it built some kind of weird Javascript-driven static site instead of just the cross-linked pure HTML thing I would have built by hand. It took me about an hour (although I did maybe ten minutes of actual work). Umami, hosted on PikaPods. For my blog, I do also pay for Netlify analytics because I find JS-based analytics misses >50% of technical users, but for stuff like this Umami is fine. The one-sentence explanation for why is that AI models already randomly select from a handful of top tokens, and watermarking just replaces that random choice with a bias that is predictable while still being equivalently “random”: as a simple example, instead of “pick randomly from the top three tokens”, you could do “count the letters in the previous ten tokens, take mod three, then pick that token”. ↩ Some notes from the vibing: GPT-5.6-Sol put extraneous text all over the page I had to get it to remove, it chose the now-very-recognizable styling that I had to rip out, and it built some kind of weird Javascript-driven static site instead of just the cross-linked pure HTML thing I would have built by hand. It took me about an hour (although I did maybe ten minutes of actual work). ↩ Umami, hosted on PikaPods. For my blog, I do also pay for Netlify analytics because I find JS-based analytics misses >50% of technical users, but for stuff like this Umami is fine. ↩

0 views
alikhil 3 days ago

How to not burnout

As someone who has experienced burnout, and has talked to many people who have experienced it, I can tell you that it’s tough to recover from burnout. It may take a lot of time and money and can cost you your job or even your profession. However, it’s much better to prevent it by taking some precautions and following simple rules. It requires less effort, helps prevent burnout and costs less. You won’t find a secret recipe or a silver bullet here. I genuinely believe that burnout can be prevented by following a few simple rules, or better yet, by treating them as hygiene. Doing the same thing every day can become extremely boring. It’ll slowly kill your motivation and you start hating your job. Automate repetitive work . Reduce the maintenance burden. For example, if you find yourself handling the same type of ticket over and over, build self-service workflows, or at least write instructions so other engineers can handle it without bothering you or your team. After you spend enough time repeating the same action many times, you’ll likely end up being good at it. Your colleagues will notice that and you’ll be asked to do this action even more. Boredom is another risk of repetitive tasks. If you find yourself bored with the same work, problem, technology or anything else, find a new challenge: something new to discover and solve, a new problem, a new tool or skill to learn. Remote work, lack of verbal communication, and absence of feedback could lead to loss of sense of reality, causing you to start doubting our skills, competence, and impact. Ask your manager and your peers for honest feedback . Accept it. Your self-doubt will likely disappear and you’ll get some direction on how to improve your skills and grow. Being motivated and productive at work will lead to big results, doubtlessly. However, for most people, life is not limited only by work. Chronic overworking will damage other areas of your life. While your productivity could benefit in the short term, in the long term you’ll suffer from tiredness, lack of motivation, and bad sleep. So try not to overwork when you can . If you had to overwork because of a strict deadline or an urgent delivery, try to compensate it by taking days off soon afterward to recover. Don’t check your Slack/email after working hours . You need a real pause. Reading work-related messages will keep your brain focused on solving problems instead of resting and being present with your friends and family. You are still working, and, in fact, overworking. If you can, I’d recommend having a separate smartphone for work-related things. You may say: “But if I don’t check Slack/email, I might miss something important and urgent.” I believe this concern is overestimated. Most communication can wait until the next morning. For really urgent matters, on-call practices should be in place, and another tool should be used, like PagerDuty. Use your paid time off days . I’ve met people who don’t use their PTO days; they stack up until they vanish. Disconnect from your work, get new ideas, travel and catch up with other areas of your life. Visit friends and family. It’s a basic necessity. It’s not surprising that minimum vacation days are required by law in many countries. Get regular physical activity. I’m not saying you should master volleyball or run 10 km every day. It could be anything convenient for you, like swimming in the sea, doing yoga, walking 10,000 steps, hiking – whatever suits you best. It can help to relieve stress and clear your mind. It’s especially useful to do it just after you finish your working day, to switch contexts. Physical activity can be a smooth bridge from your work back to your life. Have a hobby. Find something you’ll enjoy doing. It’s better if your hobby has nothing to do with your job – ideally, it should be something totally different. For example, I spend all day sitting at home in front of my computer as a software engineer. So as a hobby I’d choose something I’ll do away from home and with other people. Like going to improv comedy classes, or playing padel, or playing board games. Imagine loosing your job or burning out and when work is the main part of your self-identity. The crisis will hit you hard. If you have no idea what could be your hobby, go and try new things. Ask your friends about their hobbies, ask whether you can join them. Look for activities online, on platforms like Meetup. Think of things you enjoyed doing in your childhood. It’s easy to fall into workaholism, start overworking, forget to recharge and neglect other aspects of your life. Keeping work and life balanced takes effort. Find what works for you. What do you do / don't do to not burn out at work? Tell the world!

0 views