Latest Posts (20 found)

I added a blogroll

I realized that it might be nice if you happen to stumble on this website if I had a way to recommend other websites you might enjoy. As it turns out this is a "blogroll", a concept I have never heard of before today but whatever. It's still a good idea. I tried to add some CSS to make it easier to follow and search, but let me know if you think I missed a great site that people should check out. I'm also always on the hunt for more good stuff to read. You can find my email and social on the About page above. Also if you want to just take this Ghost theme and use it yourself feel free: https://gitlab.com/matdevdug/minimal-ghost-theme Anyway here is my new blogroll: https://matduggan.com/blogroll/

0 views
Unsung Today

“Try quickly typing 1+2+3. I bet you won’t get 6.”

Earlier this month, I talked about a rotation button in photos that behaved really nicely in iOS, and not so great on the Nothing Phone . Here’s a story of a similar fumble iOS once made that might bring the point home even more. The calculator app has been preinstalled on iPhones ever since their debut in 2007. For the longest time it hasn’t been anything more than a standard four-function calculator with a decades-old feature set. If you’re not careful, however, you can mess up even that. Ten years into iPhone’s history, iOS 11 introduced a problem just like the Nothing Phone rotation – quickly tapping on keys would show them as responding, but the actual action wouldn’t be registered. Michael Tsai’s aggregator’s first entry has a video from Stephen Heaps: It shows typing 1+2+3+4 where iOS forgets one press of +, resulting in 1+23+4 = 28. Many more people posted about it afterwards, and showed various other examples . It is oddly enthralling to see a computer fail at basic math. But what’s particularly historically interesting and perhaps even more embarrassing for Apple is the absolutely rich history of solving this kind of a problem. Calculators evolved alongside typewriters as the earliest devices with button-like (as opposed to piano-like) keyboards. But the stakes were different. Imagine a badly constructed typewriter and all the ways it can disappoint you: the letter might be faint if you press the key lightly or puncture the paper if you press it too hard, the output might be misaligned, or the typebars will jam in some way, forcing you to go again. A typewriter has to work hard to divvy up a blank, analog piece of paper into a reliable grid via escapements, ratchets, and so on. But a calculator’s work to convince the analog world to be digital is more important. After all, it’s not likely that the typewriter key you pressed will output the wrong letter – but on a badly constructed calculator, a light press of 5 could absolutely output 4, or 6, or 4.5. And, while the typewriters only take your words verbatim, the calculator’s job is precisely to create new numbers out of the numbers you type. An imprecise mechanism can mess up that math. A jam could perform a partial or nondeterministic calculation. Adding 1 to 999,999 and the force necessary for the resulting cascading carry could break a device in the middle of work. On top of all that, languages have a built-in redundancy. Evn if yuo mak many typoes, th sentece can stil be understod. But all numbers basically look alike. A calculator could make a mistake when it comes to a number that is absolutely vital for your payroll, for engineering, or for navigation – and you would never spot it. Understanding all this, many calculator makers even already in the 19th century spent a wild amount of effort convincing people not just that their devices are helpful, and fast, and easy to use, but also that they can be trusted . Buttons were carefully weighted. Comptometers came with a locking mechanism. If a machine felt something didn’t go right, it would stop working and require a hard reset. The message was: “You can trust me, because I won’t ever show you bad math, and I’ll stop myself before I will ever lie to you.” Charles Babbage was so confident in his Difference Engine that he welcomed people to try to mess with its mechanical wheels in the middle of the calculation, confident even a sabotaged machine won’t ever make a mistake. Just like with the Selectric decades later , those things were solved by people who cared, in the much harsher mechanical conditions. Of course, I don’t expect everybody at Apple core iOS team to be a calculator UI historian (although it would be nice for at least one person on the team to be one!). It is embarrassing that no one on the team had enough imagination to realize that making a button respond to a quick press during animation, but not register it would cause all sorts of serious trouble. (The bug was fixed in iOS 11.2 by removing the animations, and subsequently the animations were brought back without the original problem in iOS 11.3.) But maybe the bigger embarrassment is that Apple didn’t have a battery of tests to run on top of the UI at various speeds, mimicking fingers of what must be millions of people using the calculator app. That, too, has been a standard procedure for decades. = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/try-quickly-typing-123-i-bet-you-wont-get-6/2.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/try-quickly-typing-123-i-bet-you-wont-get-6/2.1600w.avif" type="image/avif"> = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/try-quickly-typing-123-i-bet-you-wont-get-6/3.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/try-quickly-typing-123-i-bet-you-wont-get-6/3.1600w.avif" type="image/avif"> Those tests seemed missing in 2017. I hope 2+3+4 years later that’s no longer the case. #apple #bugs #flow #history #real world #touch

0 views

Apple iPhone Air

tl;dr: I bought an Apple iPhone Air , not because I believe in Apple ’s self-congratulatory “privacy is a fundamental human right” marketing, but because after years of replacing Google Pixel hardware roughly every 18 months I have, again, reached the point where I need exactly one mobile device that will reliably run the closed-source financial, corporate and government apps that modern societies insist on. The Air is a fascinating piece of engineering, despite the new Liquid Glass in iOS being the UI equivalent of a bad car accident , and the device’s raw performance for things like mobile RAW photo development is, sadly, significantly ahead of anything Android offers in 2026. I still do not recommend using an iPhone as a primary device, because Apple remains a surveillance company. Long-time readers of this blog will know that I have spent the better part of the last decade trying to engineer my way out of the surveillance economy. I switched away from iOS and stock Android to GrapheneOS a long time ago, I run my own infrastructure , my own cloud , my own mail , my own Lastpass ( update ), my own Dropbox , my own search , my own Spotify , my own messenger , heck, even my own Zapier (sort of), and have generally tried to minimize the surface area that any single vendor has on me. The fact that I just spent a frankly absurd amount of money on a brand-new Apple phone should therefore probably come with an explanation, which is what this post is. It’s also a review of the iPhone Air itself, and a brief tour of how badly iOS has aged on a device that costs roughly the same as a decent used car in many parts of the world. Disclosure: I paid for the device with my own money. I have no relationship with Apple , and would honestly prefer not to give them my money in the first place. The opinions in this post are entirely my own. To understand why a GrapheneOS apologist would willingly buy a brand-new iPhone , it helps to know the road that led me here. The tl;dr is that for the past four years I have run a compartmentalized multi-phone setup , with one primary device on GrapheneOS that is as clean of corporate spyware as it gets, and one dedicated spyware phone that runs all the banking, travel, government identity, and other apps that I have to use to participate in modern society , but that I refuse to have anywhere near my actual life. The iPhone Air is the latest incarnation of that second device. Before the Air I kept an iPhone 11 Pro Max around for almost seven years. Of every smartphone I have ever owned across the past two decades, the only two devices that survived such long periods of time without a single hardware-level fault were the Motorola Razr V3i (yeah, I’m that old) and the iPhone 11 Pro Max . And with most of the banking, government, and corporate know-your-customer software that I am forced to interact with on a regular basis simply refusing to run on a de-Googled device, and, more importantly, being a gigantic PITA to recover onto any new phone if the current one should ever give up, it made sense to keep one rock solid device around that I could rely on. And because iOS is the best-supported target for specifically these apps, the iPhone is the only realistic alternative for someone who refuses to deal with Google Play Services . Running this kind of software on a separate physical device, that is kept in a Faraday bag when not in active use, is the cleanest way I have found to limit the blast radius of surveillance capitalism. In other words, the iPhone 11 Pro Max was not my phone , but more like my spyware appliance . The decision to specifically go with the iPhone (17) Air , as opposed to a more mainstream iPhone 17 or even a used predecessor, came down to two things: form factor, and longevity. I wanted a device that wouldn’t add a ton of weight to my travel setup , yet was still capable enough to survive (ideally) the next decade. I intentionally didn’t buy a used iPhone , because I’m planning for this device to hold up for a very long time, so I wanted Apple ’s initial one-year warranty, plus the reassurance that no previous owner had beaten up the battery, burned in the screen or let the device melt or freeze in their car for hours. For a spyware phone that occasionally rides along in a pocket as a backup, or that travels along in my carry-on, I do not want a brick. The iPhone 11 Pro Max that I am retiring weighed a frankly miserable 226 grams and the current iPhone 17 Pro Max is even heavier at 233g. The non- Max variant is only slightly lighter at 206g, and even the regular iPhone 17 , despite having a .2" smaller screen, is still 12g heavier than the Air . The iPhone Air , by contrast, is 5.64mm thin (excluding the camera plateau) and weighs only 165 grams, all while carrying a 6.5" OLED display. For comparison, the Google Pixel 6a , which for some time had been my spyware device and that the Air is also replacing, is a 6.1" device made predominantly of cheap plastic and glass, and weighs 178g. That’s 13g more than the larger and sturdier iPhone . The Air is, by far, the lightest 6.5" flagship I have ever held. While its footprint is significantly larger than I had expected from photos, and it is definitely not a one-handed device for people with small- to medium-sized hands, if you, like myself, are coming from a Pro Max , you will probably feel like you are holding a piece of cardboard, and I mean that in a positive sense. With most of the mobile phone industry having converged on the same slab-of-glass-with-a-camera-bump template, and the differentiators usually being marketing language rather than actual engineering, the iPhone Air is somewhat of an exception. The frame is Ti-6Al-4V , which is supposedly made with 80% recycled titanium . And while Apple has used titanium on Pro -line iPhones since, I believe, the 15 Pro , the Air clearly pushes the thickness budget to a different level. The closest competitor in the realm of well-established premium smartphones, which I believe is the Samsung Galaxy S25 Edge , is slightly thicker and is shown in many reviews to have build-quality compromises. Because the Air is too thin to fit a conventional stamped USB-C connector, Apple ’s engineers 3D-printed the metal connector frame out of titanium powder, fused with a laser, and machined to spec. As reported by Engadget , Apple ’s justification is that at this thickness, there is no other way to fit a standard-compliant USB-C connector. I am skeptical of the claim that this was the only option (a redesigned stamped part would presumably also work), but the fact that they shipped a 3D-printed titanium structural component in a consumer phone is an interesting engineering choice. Another interesting engineering choice is the camera plateau . Rather than the widespread camera trypophobia bumps , the Air has a raised horizontal rail at the top of the back, under which sit the A19 Pro SoC , the new C1X modem, the networking silicon, and the single 48MP Fusion camera. This concentration of components in a thicker strip is what makes the rest of the chassis flat. From a thermal perspective it makes total sense, as it concentrates the device’s heat generation in a region that is the easiest to keep away from the user’s hand. This is one thing that always bugged me about the 11 Pro Max , which is that the moment the device is under load (which is pretty much all the time when running a recent iOS version) the backside area, where your fingers naturally rest when holding the phone, will get uncomfortably hot. On the Air you need to actively move your fingers underneath the plateau to feel the heat. However, because the Air does not have the vapor chamber that the Pro line introduced this generation, it throttles more aggressively under sustained load than the 17 Pro . For my use case (banking apps, occasional photo editing, very occasional video) this is irrelevant, but for someone trying to play a demanding game for two hours straight, it probably isn’t. Speaking of heat, the Air packs quite a lot of it with its 6-core CPU with 5-core GPU and the Neural Accelerators , which in benchmarks scores 9,497 points in Geekbench 6 multi-core. For comparison, the Pixel 10 Pro ’s Tensor G5 , in the same benchmark, sits at roughly two thirds of that number . The C1X modem, which is Apple ’s second-generation in-house cellular modem, is also an interesting piece of engineering. It has replaced the Qualcomm silicon that has been in every iPhone since the 12 and in real-world testing the C1X appears to be noticeably more power-efficient than both the Qualcomm modem in the 16 Pro and the Samsung Exynos 5300 modem that has been making the Pixel family miserable for the past several generations. For a device whose battery is a comparatively modest 3,149 mAh , that efficiency gain is what makes the Air ’s battery life usable at all. One big difference from literally every phone that I have ever owned is the fact that the Air has no physical SIM tray. If I were planning to use this phone as my primary phone, this would be a huge PITA for travel . Pre-paid SIMs in many countries are usually not available as eSIM, or when they are, they’re significantly more expensive and privacy-invasive due to KYC measures. I have written before about why I strongly prefer physical SIMs for both privacy and practical-travel reasons, and Apple ’s decision here is a real downside for any full-time traveller . The reason for the lack of a physical SIM tray appears to be that Apple needed the volume for the already small enough battery, which makes sense. It is nevertheless a regression. Another engineering decision that somewhat made me question the whole thing is the USB port. In 2026 the world’s most valuable company decided that, on a flagship-priced device, you still only get USB 2.0 transfer speeds out of the USB-C port. For whatever reason, Apple reserves USB 3 speeds for the Pro line. This means that transferring 50MB RAW files off an SD card via a USB-C card reader is not exactly slow , but it is also not fast , and it is straightforwardly insulting on a device as expensive as the iPhone Air . Plus, the whole USB design choice also leads to absurd incompatibility issues with plenty of built-in USB controllers within SSDs, making a good number of external drives simply unusable with the Air . One big letdown at this price point for many people seems to be the single 48MP rear camera. Coming from an iPhone 11 Pro Max , or even the cheaper Pixel 6a , both of which had multi-camera arrays, I understand that the average user might perceive this as a step down for photography. However, personally, I don’t care too much as I take essentially all of my photos on dedicated cameras anyway. GSMArena ’s review calls the single rear camera the device’s “big pain point” , and I think that is fair at this price. Taken together, the Air is probably the most opinionated phone Apple has shipped in a long time. The company was willing to sacrifice camera count, port performance, and thermal headroom to deliver a piece of hardware that is as thin and light as it is, and that is built around in-house silicon. Whether this approach turns out to be a good idea, history will tell. As a piece of standalone consumer electronics, however, it is pretty impressive, and I say that as someone who, in almost every other domain , finds modern consumer hardware profoundly depressing . And then there is the software. iOS 26 ships with Liquid Glass , Apple ’s new design language, which applies a translucent, light-refracting, glass-like aesthetic to basically every system surface, from the lock screen to the notifications, Control Center, app icons, menus, and even system alerts. Apple ’s clearly high af marketing department describes it as “a new material that combines the optical qualities of glass with a fluid, responsive feel that brings depth and dynamism to every interaction” … whatever that is supposed to mean. What it actually feels like, in daily use, is sadly less flowery , and the user backlash and performance issues have been covered extensively . The kindest thing I can say about Liquid Glass is that, despite it being the Ferrari Luce of UIs , the engineering underneath it is impressive. The amount of real-time blur, refraction, and material simulation Apple is pulling off on the A19 Pro at 120Hz is a graphics feat that the Android system can only dream about. Having said that, however, I don’t think that any of that engineering would be truly necessary if somebody at Apple had asked the question “does this design language actually make the device easier to use?” before shipping it on a billion devices. The primary purpose of the Air for me, as with its predecessors, is to host all of the closed-source corporate spyware that I refuse to put on my primary GrapheneOS device . This includes: On the Air , all of this just works . Apps launch instantly, attestation succeeds, and the device does not get anywhere near as hot as the average Pixel phone doing any of it. After four years of fighting with Google ’s hardware and Android ’s gradual decay, I am, embarrassingly, enjoying the boring reliability of iOS . The thing I did not expect, and the thing that has surprised me about the Air , is how good it is at mobile RAW photo development. For a long time I had been doing all of my photography workflow on a Pixel Tablet running GrapheneOS , using Lightroom Mobile . The arrangement worked, but it didn’t work well. Lightroom Mobile on the Tensor G2 -powered Pixel Tablet has increasingly become a sluggish, crash-prone mess , with editing operations that take noticeable seconds to render preview updates, occasional import failures, and a battery life under heavy editing that is less than three hours. On the iPhone Air , the same workflow (using a USB-C SD card reader, plus the free Snapseed app) is vastly faster. Importing a card full of 50MB RAW files from my Fuji is bottlenecked by the USB 2.0 port rather than the phone. Edit operations on RAW files are essentially instantaneous, even with multiple layered adjustments. Exporting to JPEG happens in well under two seconds per image. Obviously the Air ’s A19 Pro SoC was released several years after the Tensor G2 , and it is orders of magnitude faster on graphics workloads, but the Pixel Tablet is nevertheless a 2023 product that Google is still selling new in 2026. The fact that an in-pocket, 5.6mm-thin iPhone runs circles around it on the same workload is, frankly, embarrassing for Google . The mobile content-creation software ecosystem on iOS is also significantly better than what is available on Android , and not just because the hardware is faster. Lightroom Mobile on iOS is more polished than the Android version, and even Snapseed , an app by Google , runs noticeably more smoothly on iOS than it does on Android . Even free, ad-supported photo apps tend to be visibly better on iOS than anything on Android . Some of this is because developers prioritize iOS for revenue reasons, because literally every little sh.t app on iOS these days can seemingly charge a price or, worse, a subscription fee. However, some of it is also because Metal and Core Image are better-integrated graphics and imaging stacks than what Android exposes to third-party developers. Whatever the reason, in the specific category of mobile RAW photo development, iOS is not just ahead, but it is literally pulling Android ’s pants down and slapping its Baklava . Even setting RAW editing aside, the device feels noticeably snappier than any Android phone I have used, including the Pixel 8 . App launches are faster, animations are smoother, and background apps stay in memory longer. Apple ’s memory management on iOS , even with comparatively modest RAM (the Air has 12GB), is more aggressive about keeping apps warm than Android ’s. Tom’s Guide ’s real-world performance testing shows iPhones consistently lapping Android flagships on workloads like video transcoding, where the iPhone 17 Pro models complete a standard test in 22 seconds, while every Samsung Galaxy S25 variant takes over twice as long. On purely synthetic Geekbench 6 multi-core, the Snapdragon 8 Elite Gen 5 does outperform the A19 Pro , but the Snapdragon ’s lead in synthetic CPU benchmarks does not translate consistently into real-world responsiveness, because Android as an operating system still carries a significant amount of overhead that even the fastest silicon cannot entirely paper over. On the iOS side, the A19 Pro ’s single-core performance (around 3,871 points in Geekbench 6 ) is still the highest of any shipping mobile SoC. A heuristic that I have used for years, and that has consistently held up, is that an iPhone from year N tends to feel as snappy in everyday use as an Android flagship from year N+1 or N+2 , and I’m certain that the iPhone Air is no exception to that rule. This is a generic OS observation rather than a comment on the Air specifically, but it is a real observation, and it is one of the reasons why iOS devices keep getting recommended even by people who would much rather be using something else. Besides the superior hardware of Apple ’s phones over specifically the Google Pixel devices, the ecosystem of accessories that are intentionally designed to fit the iPhone lineup is another thing that Google , as well as most other Android manufacturers (with the exception of Samsung and Xiaomi ), has failed to establish. In recent years, Apple managed to further increase their ecosystem’s lead with their MagSafe quick-attach feature, for which you can find literally everything, from snap-on card holders through powerbanks all the way to actual stands. Because I was curious to give the MagSafe system a try, I decided to pick up a powerbank that unintentionally fits the “Space Black” iPhone Air better than Apple ’s own, plain white iPhone Air MagSafe Battery . I chose the Xiaomi UltraThin Magnetic Power Bank 5000 not only because its design was clearly targeted at iPhone Air users, but also because it is the thinnest (6mm) and lightest (98g) 5000 mAh power bank on the market, thanks to its relatively new high-energy-density silicon-carbon battery. When attached to the back of the Air it brings up the device’s weight to 263g in total, which is noticeably heavier than most other smartphones, but it also almost doubles the phone’s battery life. While the Xiaomi power bank can deliver 22.5W, it only does so via USB-C, and because it is not MagSafe -certified, it drops to 7.5W when charging the phone wirelessly (by snapping it on using MagSafe ), rather than the 15W that Qi2 phones get out of it. With the Air locked and inactive it takes about two to three hours to charge it to around 92%. It never reaches 100%, despite the power bank carrying considerably more capacity than the phone’s integrated battery, due to the inefficiency of wireless charging . Note: I’m aware that Google introduced MagSafe compatibility into its Pixel line with the Google Pixel 10 , however, there is no dedicated ecosystem targeting specifically the Pixel and its hardware design . For the narrow use case that I have personally been struggling to solve, which is “I need a single, lightweight, and highly reliable device to host all the closed-source corporate spyware I am forced to interact with, and I already have a primary GrapheneOS device for my actual life” , the iPhone Air seems like a very good option, at least ignoring the absurd price tag. It is light enough to carry in a Faraday bag inside my cabin luggage with me and it’s flat enough to occasionally bring it alongside a Pixel without making my pockets look like I pack big data. It appears to be reliable enough that I do not expect to repeat this exercise for at least another five to seven years, and it runs every attestation-dependent banking, travel, and near-future dystopian government app I might need it to run. On top of that, the mobile content-creation software story on iOS is, sadly, lightyears ahead of Android , as well as Windows and Linux . For anybody whose use case is “I want this device to be the primary, always-with-me phone that handles my nudes photography, my manifestos notes, my 4chaning web-browsing, and my doomscrolling actual life” , I cannot in good conscience recommend the Air , or any other iPhone , for exactly the same reasons I have been writing about for the past years . Apple is, with all due respect to its marketing department, a surveillance company . The fact that it chooses to surveil its users somewhat more politely than Google does is not, by itself, a reason to grant it custody of your private life. You cannot audit the OS, you cannot disable any hidden surveillance subsystem, and you most certainly cannot install a hardened replacement. The iPhone Air is a fascinating piece of engineering and a perfectly competent spyware appliance . It is however not, and will never be, my primary phone. Banking apps for several jurisdictions, most of which now require Google Play Integrity or equivalent attestation on Android and basically refuse to function on GrapheneOS without disabling the very protections that make GrapheneOS worth running in the first place. Apple Wallet for boarding passes and, in some jurisdictions, virtual VISA and MasterCards. Luckily I’m neither a citizen nor a resident of a country that imposes the absurdity of digital state-issued identifiers . Corporate apps that some clients might use and require, as well as privacy-invasive bs like WhatsApp , using throwaway phone numbers, in jurisdictions in which large parts of public life run on it . Travel apps (airlines, hotels, rental cars), most of which are technically available on GrapheneOS but which are also ill-behaved, ad-laden, and absolutely should not be on the same device as my personal data.

0 views

📝 2026-07-20 08:10: Yup. Goat-proofing the chicken door really worked a treat. FML. 🤦🏼‍♂️

Yup. Goat-proofing the chicken door really worked a treat. FML. 🤦🏼‍♂️ Thanks for reading this post via RSS. RSS is ace, and so are you. ❤️ You can reply to this post by email , or leave a comment .

0 views

Thoughts on “To Have or to Be” by Erich Fromm

As I mentioned in my IndieWeb Book Club post , I first read Fromm’s book more than 20 years ago. I was in my early high school years when I stumbled upon his work and was immediately intrigued by his way of looking at the world. The fact that he was not “just” a philosopher but also a psychologist and a sociologist meant his ideas felt more concrete and applicable in today’s world, even though his book was more than 30 years old at the time and is half a century old now. In reading To Have or to Be a second time, I was surprised by a couple of things. I remembered quite well the core thesis of the book, this tension between the two fundamental modes of existence, one predicated on the idea of possessing things, while the other is built on the concept of living and experiencing things without feeling the need to have them. It’s a way of looking at the world that felt right in my youth, and it still feels that way now, probably even more so considering the state of the world. But what I didn’t remember was the second half of the book, where Fromm tries to outline what a new human being and a new society could look like if we ditched the first mode—which is basically the pure capitalistic and consumeristic mindset—and embraced the second. He was certainly very optimistic, and there are passages that made me smile: seeing how convinced he was that the capitalist model was on the cusp of breaking down is heartwarming. But at the same time, there are thoughts in that second part that are shockingly accurate in describing today’s society. His thoughts about news, about the manipulation of people through advertising and propaganda, I was quite surprised to see how well they aged. Reading the book again made me want to pick up a bunch of his other books I also read in my teens, from Escape from Freedom to The Anatomy of Human Destructiveness , but also made me want to read some of his other books— The Art of Being seems very interesting—and also made me want to read Marx and Spinoza. I might add a bunch of those books to my reading list. Overall, very happy to have read the book a second time. I’m not a fan of re-reading books, but this was definitely worth it. Thank you for keeping RSS alive. You're awesome. Connect via email :: Sign my guestbook :: Support for 1$/month

0 views
Unsung Today

“Is this acceptable? Not really. Is it understandable? Absolutely.”

When you visit parts of the website of the Driver & Vehicle Licensing Agency in the U.K. outside of standard office hours, you will see this: = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/is-this-acceptable-not-really-is-it-understandable-absolutely/1.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/is-this-acceptable-not-really-is-it-understandable-absolutely/1.1600w.avif" type="image/avif"> It seems ridiculous, almost like the famous 500-mile email bug , but Dafydd Vaughan explains how it came to be . It won’t be a surprise that the story starts with an old computer system, built around a day/night cadence where people would visit the agency during the day, the information would be collected, and all the data processing would happen overnight: At the time, many of DVLA’s services – particularly those relating to driving licences were still backed by an old IBM mainframe from the 1980s – fondly known as Drivers-90 (or D90 for short). D90 was your typical mainframe – code written in COBOL using the ADABAS database package. Most data processing happened “offline” – through batch jobs which ran during an overnight window. The team upgrading it in 2013 decided to not attack too many things at once, and swap the innards but keep the day/night system intact. I’ll let you read some details inside if you are interested, and if you want to learn why the 2013 decision is still a 2026 (or at least 2025) reality. Like many of these , it is a story of heavy financial and technical constraints. But I will excerpt this bit, because I think many teams in any company – design systems, security, or “core” (whatever that happens to be) – would relate to it: It’s difficult for an organisation to keep its focus and attention on a complex upgrade – particularly without getting noticeable benefits along the way. The tricky part is that in many places you do end up in a legacy situation that is exactly as absurd as “an internet service having opening hours,” but you might be too close to the problem to even realize it. #maintenance #software evolution

0 views
Dan Moore! Yesterday

What It’s Like To Lead An Engineering Org: Thoughts On “CTO In The Loop”

I just read CTO In The Loop by Balki Kodarapu. It’s free on Kindle through July 20th, so if you want to give it a read, go do that now. It’s a narrative arc following a developer, “Sam”, who starts as a software engineer at a startup, and ends up a fractional CTO after a career that includes engineering manager, director of engineering and VP of engineering at a company that IPOs. The book was a fun read, and does a good job describing what it’s like to build a company and a software business from the early days to a more mature organization. The realities of Sam’s growth ring true: he starts out focused on features and clarity, and then moves up levels of abstraction. He loses touch with tasks and areas that used to matter, but gains higher level views and more ability to influence the organization. As a director, if you’re still reviewing PRs the way you used to, you’re basically breaking your organization We also see the evolution of several other characters, including Priya the project manager, Mei the UX designer and Jack the senior developer. Some of these folks stay with Sam across companies, while others pop in and out, just like real world colleagues. Some details felt a little forced; for example, locking down a GitHub repo and preventing merges to the main branch are treated as a major effort, when most software shops do that pretty early. I especially liked the portrayal of the thorny moments of management that Sam encounters: I also really enjoyed the emphasis on integrity, and how the protagonist was repeatedly challenged to hold onto it through different stages of business growth. I think integrity is an easy thing to write about and much harder to actually live. In the moment, you tell yourself that just this one tweak that you might feel uneasy about will move the business in a positive direction. But doing this repeatedly only works in the short term. The benefit is immediate and the cost is deferred, and avoiding these kinds of decisions is a constant challenge as an organization grows. The hard part is that it’s rarely obvious where the line is. Most people agree you shouldn’t commit fraud. But there are dark UX patterns that benefit the business. For example, you could make it much harder to unsubscribe from a service than it was to subscribe; every organizational incentive points that direction. Another example is the integrity in making it easy for a customer to leave: making it easy for people to export their data and take it with them. That’s a frustrating tradeoff for whoever’s job it is to sell the product, but it’s a pretty clear-cut example of integrity. You want to earn people’s business, not hold their data hostage. There are other places integrity shows up over the course of a career, but the author does a good job of showing how it is hard to maintain, especially when you are focused on the day to day of building a software business. CTO In the Loop is a fun, easy read. I read it on the Kindle app in about 2 days while on a vacation. I’d highly recommend it to people early or mid-career in software who are curious what life looks like at a different level of the org chart. When I was younger, I assumed people at higher levels had a lot more control, agency, and power. That’s true, but they are also further from the customer, removed from knowledge of how technical systems actually work, and have a lot less clarity about the day to day. To say nothing of many competing demands that folks at the director or higher level have on them. Writing code, or telling an AI how to write code, is one thing. Thinking about what the company looks like in one year or five, how to bring different teams together, how to manage investors and understand the market, and how to keep every department rowing in the same direction: that’s a tremendously difficult job. This book is an accessible narrative of those tensions, to which I don’t think new or even mid-career developers usually get exposure. So go check it out. Did I mention it’s free on Kindle through July 20th? Balki, thanks for writing this excellent, interesting fable of a software developer’s career. firing a team member multiple rounds of layoffs (and subsequent hiring) changing his understanding of the constraints and goals of the engineering org he is in

0 views
Kev Quirk Yesterday

📝 2026-07-19 19:09: Today was spent goat-proofing the automatic chicken door. We recently moved it to ground level...

Today was spent goat-proofing the automatic chicken door. We recently moved it to ground level and now the goats use it as a scratch post. 🐐 Thanks for reading this post via RSS. RSS is ace, and so are you. ❤️ You can reply to this post by email , or leave a comment .

0 views
Unsung Yesterday

“Maze of pages, redirects, and confirmation emails”

Nikita Prokopov on his fun Tumblr-like site Grumpy Website that collects UI transgressions with light commentary: = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/maze-of-pages-redirects-and-confirmation-emails/1.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/maze-of-pages-redirects-and-confirmation-emails/1.1600w.avif" type="image/avif"> Why log in when you have to, when you can instead experience the same Kafkaesque maze of pages, redirects, confirmation emails, copying codes, and searching for a second device—all a week in advance? I agree with authentication experiences being absolutely horrible, especially in the corporate context. But I think the “your login expires in 5 days” callout is actually positive. Your reauthentication will happen anyway according to the security regimen. Instead of it potentially surprising you at a wrong moment – for example just before, or in the middle of the presentation in front of a boss – you could get ahead on that process during the next moment of downtime, and have some peace of mind. I know Slack has a similar warning, although I don’t happen to have a screenshot handy; I imagine a few other places do, too. (However, that gesture doesn’t give apps permission to handle expiring credentials badly. There should be no excuse for losing anything the user is doing – for example writing a long comment – just because the reauthentication kicks in in the middle of them doing that. Even the app losing its mind for a while just as its separate parts are waking up to the reality of reauthentication at their own pace and throwing strange errors, but the login screen hasn’t reappeared yet, can be really confusing.) #attention #nikita prokopov #security

0 views
Kev Quirk Yesterday

I've Moved Back to GitHub

Back in May, I wrote a post about how I'd migrated all my repos away from GitHub. This ended up being a combination of my self-hosted Git server on my Synology (for my private repos) and Codeberg for my public ones. Well, since then Codeberg has just given me problem after problem. I had issues working with repos via SSH for weeks . It just kept timing out, or taking an age to actually do anything. So I switched to HTTPS instead, which seemed a little better, albeit still generally very slow. This morning I came to do a bit of work on some issues and PRs that have been logged against the Pure Comments repo and (unsurprisingly at this point), I was greeted with this: I've been trying to get to the repo for around 30 minutes now, and it just keeps failing. I can't work like this - I'm a busy guy, so when I do get time to work on these fun side projects, my shit needs to work. People give GitHub a hard time for outages, scraping, etc. but in all the years I've used it, I've never had an issue. It has always been quick, and it always worked. In my boredom while waiting for Codeberg to sort themselves out, I decided to peruse my RSS feeds and I came across this post by Sal . It seems he's been having similar problems with Gitlab. This was the final straw. I've had enough of battling with Codeberg, so I've switched back to GitHub. Some may not like this decision as Codeberg is generally considered the more community friends code hub, but I need to make a pragmatic choice here as the constant issues are sapping most the fun out of these projects. I'll continue to host my private projects on the Synology, as that works great. But as of right now, we're back on GitHub for all [my projects . I think I've caught all references to Codeberg across the various sites and docs (thanks to Gemini), but if you see a problem, please log an issue. There is 1 final update on Codeberg for both Pure Blog and Pure Comments. This points the updater back to GitHub, from then on, subsequent releases will be on GitHub. I'm bound to get some comments and suggestions about other options other than moving back to GitHub, so I'll try and hit them before you take the time to comment/email with recommendations: What about a self-hosted Forgejo instance? Absolutely not. I don't have time to maintain something like that. I'd rather spend my free time working on fun projects than managing infrastructure. But GitHub are really bad because of what about ? Nope, sorry. I don't want to use other platforms that potentially introduce the same issues I've been having with Codeberg. Look at Sal's post further up - he was on Gitlab and having similar problems to me. This is disappointing, are you sure about your decision? I'm very sure. It is disappointing, I agree. But I've given it a lot of thought over the last few weeks and I'd rather go with the tool that I know will work, than the one I have to battle with. Plus, this is all public code, so I don't think I'm giving anything away by switching back. I'll NEVER use GitHub. I'm gonna stop using now. That's a shame, but it's your call. Thanks for reading this post via RSS. RSS is ace, and so are you. ❤️ You can reply to this post by email , or leave a comment . What about a self-hosted Forgejo instance? Absolutely not. I don't have time to maintain something like that. I'd rather spend my free time working on fun projects than managing infrastructure. But GitHub are really bad because of what about ? Nope, sorry. I don't want to use other platforms that potentially introduce the same issues I've been having with Codeberg. Look at Sal's post further up - he was on Gitlab and having similar problems to me. This is disappointing, are you sure about your decision? I'm very sure. It is disappointing, I agree. But I've given it a lot of thought over the last few weeks and I'd rather go with the tool that I know will work, than the one I have to battle with. Plus, this is all public code, so I don't think I'm giving anything away by switching back. I'll NEVER use GitHub. I'm gonna stop using now. That's a shame, but it's your call.

0 views
A Smart Bear Yesterday

What is the value of one hour of a startup founder's time?

Startup founders undervalue their time. Here's why you should act like it's $1000/hr, and how it changes the decisions you make.

0 views
Sean Goedecke Yesterday

Impro is a handbook for running a cult

Here’s the big idea in Keith Johnstone’s book Impro : This take doesn’t sound particularly original, but references to Impro pop up in all kinds of places: in influential tech blogs , as part of the initial process of onboarding for Palantir, and on the reading list of multiple big-tech founders . Impro is part of the secret canon of Silicon Valley, right alongside books like Seeing Like a State and The Power Broker . Why is that? For two reasons: first, because Johnstone’s outsider critique of established institutions is appealing; and second, because Impro is a handbook for running a cult. The part of Impro that is most obviously useful to software engineers is Johnstone’s chapter on status. According to him, status games pervade all social interactions. Even innocuous, friendly conversations operate in terms of status. When you apologize or downplay something to “be nice”, that’s performing low status; when you reassure somebody, that’s performing high status; when you and a friend are comparing stories, you’re making friendly bids for status from each other. In the workplace, these status games are conditioned by the formal status of your role: you must allow your boss the high status position most of the time, or you’ll be (correctly) perceived as insubordinate. This is understood in some cultures, where it’s often called “face” , but in Western cultures it’s taboo to openly discuss status games. The core social skill is the ability to deliberately alter your status. Someone who can only perform low status is a weak person, pitiable, annoying. Someone who can only perform high status is a braggart, a posturer, dangerous. To be effective socially, you must be able to switch between high and low status when appropriate, sometimes from sentence to sentence. I wrote about this exact point at the end of Big tech engineers need big egos : effective senior+ software engineers must be able to present as high status in order to be useful authorities, but also to switch to low status in order to take direction from the company leaders. As an example, Johnstone describes in detail how he manipulates status in the classroom. He begins by sitting on the floor (deliberately assuming low status), and explaining that if his students fail, it’s his fault not theirs, since he’s the expert. The initial low status puts the class at ease, but in his words, ”[my] actual status is going up, since only a very confident and experienced person would put the blame for failure on himself.” These skills are not just useful for improv comedy. Impro is not just a book about improvising well. It’s a book about how you should live your life. In other words, Johnstone thinks that everyone would be better off if they became more spontaneous and ditched their shells of over-analysis. He criticizes the culture of Western thought in a number of different areas. According to him: Johnstone didn’t come up with these ideas — they’re standard counterculture positions from the 1960s and 1970s — but it goes to show how he connected improvisational comedy to this general anti-establishment political program. Johnstone ran his classes and theatre troupe like a revolutionary cadre. Here are some quotes from Something Like a Drug: An Unauthorized Oral History of Theatresports : So of course when I was invited to join Loose Moose Theatre and train at improvisational games late at night in an abandoned garage in a run-down portion of the city, I was thrilled. I remember thinking, This is a revolutionary act. Keith [Johnstone] got a group of his more talented students together to start improvising outside of school hours. Usually in his basement. The Secret Impro group—it’s very strange. It was very much that Keith said we were going to do this, and we’d just do it. It was like we were sheep. Keith would say when we were going to do a show, and we’d just do it, blindly. Like I said, if we had the videotapes now, we’d be very embarrassed and probably never go on stage again. We became a group of people who would follow Keith. There was always that sort of “tag” put on those people who were with Keith and those people who were against Keith. We were the people, basically, that if he said something, we believed it. To some extent, it’s plausible that teaching acting or improvisation requires a high level of trust in your teacher. When Johnstone says things like “Students need a ‘guru’ who ‘gives permission’ to allow forbidden thoughts into their consciousness.”, I can believe that it’s just how you have to teach acting. But the more I read of Impro (and particularly when I read Something Like a Drug and Johnstone’s biography Keith Johnstone ), the less it sounded like an ordinary book on acting. Instead, it began to sound like a charismatic man who had found a way to gather a group of disciples that would let him mold their psyches. In other words, it began to sound like a cult . Impro was first introduced to the software world by Venkatesh Rao (of Gervais Principle fame), who wrote a brief review . Rao gives a detailed account of the first three-quarters of Impro , but glosses right over the last chapter, called “Masks and Trance”, simply saying “despite the disturbing raw material, the ideas and concepts are not particularly difficult to grasp and accept”. What ideas and concepts? Johnstone’s discussion of masks (or “Masks”, in his language — he always capitalizes the word) is as explicitly cult-like as Impro gets. In brief, Johnstone has a box of literal, physical prop masks. He introduces the box with great ceremony to his students 2 , warning them seriously about the dangers of possession and reassuring them that he is a skilled and competent spirit guide. Through various hypnosis-adjacent techniques 3 (Johnstone draws the parallel quite explicitly) he conditions his students to be in a trance state when wearing a mask, and believes this produces more authentic emotional states in their acting and improvisation. Here are some quotes from the book: A high-status person whom you accept as dominant can easily propel you into unusual states of being. You’re likely to respond to his suggestion… Once you understand that you’re no longer held responsible for your actions, then there’s no need to maintain a ‘personality’. One famous French teacher of the Mask—who won’t approve of this essay 4 —divides students immediately into those who can work Masks and those who can’t. I don’t cast an actor to play a Masked role until I know he has the ability to become ‘possessed’. It’s true that an actor can wear a Mask casually, and just pretend to be another person, but Gaskill and myself were absolutely clear that we were trying to induce trance states. Johnstone has a long and painful explanation of how new mask-wearers seem to mentally regress to the point where they don’t know how to open umbrellas or interact with chairs. He describes one student always going to the bathroom before putting on a mask, because she’s worried she might wet herself. New mask-wearers are non-verbal must be taught to speak again. If this were at the beginning of the book, I think it would turn a lot of people off. But by the time you get to it, I suspect most readers are already warmed up enough to say “sure, why not, it seems weird but I guess it works”. Not me! Johnstone attempts to defuse the obvious weirdness by arguing that trance states are very common (e.g. being lost in a book). More unconvincingly, he says this in response to the worry that vulnerable people are going to get mentally harmed: As for the fear of madness, I would answer that the ability to become possessed is a sign of correct social adjustment, and that really disturbed people censor themselves out. Either they can’t do it, or they’re afraid to even try. People who feel themselves at risk avoid situations where they feel likely to ‘go to pieces’. Does this convince anyone? Mentally vulnerable people fall into dangerous situations all the time: ayahuasca trips, cults, GPT-4o , and so on. It’s such a weak argument. In general, I’m struck by the sheer power Johnstone held over his disciples. He has them yell slurs at each other, encourages them to feel deep emotions in quick succession, relax any mental defenses and regress to a childhood state, and literally hypnotizes them . He explicitly lays out his procedure for breaking down their sense of self: The stages I try to take students through involve the realisation (1) that we struggle against our imaginations, especially when we try to be imaginative; (2) that we are not responsible for the content of our imaginations; and (3) that we are not, as we are taught to think, our ‘personalities’, but that the imagination is our true self. If your imagination is your true self, and you’re not responsible for its content, you’re not ultimately responsible for anything: you’re in the safe hands of the guru, who can mold you as he wishes. Later on, Johnstone walks it back a bit: In the end they learn how to abandon control while at the same time they exercise control. … You have to misdirect people to absolve them of responsibility. Then, much later, they become strong enough to resume the responsibility themselves. So the explicit idea is that ( much later), the guru hands autonomy back to his disciples, when they’re ready to take it. This does not exactly reassure me, particularly against the background noise of everyone in Johnstone’s circle saying “boy I sure love being part of this cult!” I don’t think Johnstone was preying on his students. The strongest evidence against this is that he did marry a student 5 , Ingrid Brind. That’s not great! On the other hand, it was fairly standard for professors back then — when I was in grad school for philosophy, several of my older male 6 professors had wives that they’d taught decades ago — so I don’t think it proves Johnstone was that kind of cult leader. I even read Ann Jellicoe’s play The Knack to get a better picture of Johnstone’s character. Jellicoe had an affair with Johnstone for several years, and his official biography claims 7 that the character of Tom in The Knack is directly based on Johnstone. The Knack is a rather unpleasant play about sexual assault, but Tom’s character is largely asexual: he’s certainly no feminist, but is much more interested in impressing people with his intelligence than with getting laid. In Something Like a Drug , two women who were part of Loose Moose, Johnstone’s Canadian improv group, describe their experiences: You know, it brings around the other question: Why do the guys get laid after the show and not the chicks? You know, I can remember those days when Tony [Totino] and Dave [Duncan] and all those guys … the women would swarm around them. Those were the days, my friend. In Loose Moose I think there are fewer women not only because of the training, but because of the guys in Loose Moose. When I came up with Joanne and Laura, there was a real initiation that was going on, and there was a group of guys at that time who were all single. And they would hit on you to the point where one night Joanne, Laura and I, who really didn’t know each other, were in a show together, started talking and realized that we were getting the same pickup lines from the same guys. And that’s when you realize what’s going on, and I think that’s intimidating. Or if a woman gets into a relationship with a senior improvisor and it doesn’t work out or something bad happens. I think that’s one reason. This dynamic doesn’t sound great, but it doesn’t mention Johnstone, and it doesn’t sound particularly unusual : I’ve heard versions of this story about all kinds of ordinary male-dominated nerd spaces. In fact, reading through the anecdotes in Something Like a Drug is a good antidote to the cultish atmosphere in Impro . Johnstone’s argument goes something like: “if we could only throw away the restrictive chains of Western culture and permit ourselves to be as obscene and free as children, we would be transported to a better, more beautiful world”. Well, you tried that, and the women in the group are still relegated to playing bimbos and housewives, there are still petty personal fights, and the guru is out here union-busting 8 . What was enlightenment supposed to look like? I think the most generous defense of Johnstone is that his group was not unusually cult-like, and that any similar account from one of his peer improv teachers would raise the same red flags. Maybe improv classes and groups (particularly in the 70s and 80s) were just cultish in general? Having now read four books on Johnstone, I’m reluctant to go and read more to prove or disprove this theory, but it’s at least plausible. To anyone familiar with San Francisco software engineering culture, it should be pretty clear why Impro is so popular. The line between a startup and a cult is very thin indeed. In his book Zero to One , Peter Thiel famously says that good startups are “slightly less extreme kinds of cults”. If you believe that, it makes total sense to assign Impro as mandatory reading for new Palantir hires. It tells them what kind of cult you’re trying to run: one where you’ll disregard existing cultural norms, learn to play status games well, think on your feet, and generally be molded by the guru into a more persuasive, more effective engineer. Read critically, Impro also serves as a handbook for engineers who are trying to recognize if the environment they’re in is cult-like. Is your company telling you to reinvent your personality in order to be better at your job? Are you under the spell of a charismatic, high-status leader? Is your company trying to keep you in an unquestioning flow trance state? In the great battle between the shackles of restrictive culture and the glorious freedom of the guru, I am always and forever on the side of the shackles of restrictive culture. In general, I think most boring and stupid social norms (such as not hypnotizing and marrying your students) serve an important purpose and shouldn’t just be cut down in the name of freedom. Impro is still a good book. There’s a lot to learn from Johnstone’s analysis of power dynamics, of education, and of creativity in general. By all accounts he was excellent at teaching students how to improvise. But I wouldn’t recommend adopting it as your life philosophy, and I’d recommend being a bit suspicious of anyone pushing this book too hard. Getting rid of the existing social structures might benefit confident, wildly charismatic gurus like Johnstone, but most of us are just ordinary animals who do better in a group governed by norms. In fairness to Johnstone, he cites Sheila Kitzinger’s The Experience of Childbirth in support of this claim (the others he just puts in his own words), so maybe he felt that this was a bit out there. As you would expect, the pain of childbirth is a universal biological fact . Concerningly, the description in Something Like a Drug (in the foreword) suggests that this class was unofficial . As an example, he prompts the masked student to relax, then startles him with a mirror to trigger the trance state. Probably Jacques Lecoq . See page 83 of Keith Johnstone: A Critical Biography . I suppose that’s redundant. On page 51 of Keith Johnstone: A Critical Biography (it’s called “critical” but it was clearly written with Johnstone’s involvement and support, and does not seriously criticize him at any point). In The Knack , Tom gives a monologue about how to teach children to play the piano that could be lifted straight out of Impro . In 1983 Johnstone “read the riot act” to the improv players who were planning to unionize, threatening that they’d be cut out of the group for good. To quote Dennis Cahill, a group member at the time who opposed the union: “I just didn’t see the point to it. … I didn’t really see a need to confront Keith or cause Keith problems or to upset him in any way over something as simple as Who Has The Power or Who Doesn’t.” Children are naturally creative, but are violently formed into repressed adults by Western culture and education The process of becoming more creative and expressive is largely a process of unlearning these habits of repression Improv — improvisational comedy — is thus not just the skeleton key for learning to act, but for unlocking a more authentically human way of life Everyone is more or less equivalently mentally ill, but “sane” people simply have better coping mechanisms Cities and “taking pills” (read: antidepressants) are obscene, but you should be able to make sexual jokes in the workplace and generally be uninhibited If we were free from the puritanical shackles of Western culture, childbirth would not be painful 1 In fairness to Johnstone, he cites Sheila Kitzinger’s The Experience of Childbirth in support of this claim (the others he just puts in his own words), so maybe he felt that this was a bit out there. As you would expect, the pain of childbirth is a universal biological fact . ↩ Concerningly, the description in Something Like a Drug (in the foreword) suggests that this class was unofficial . ↩ As an example, he prompts the masked student to relax, then startles him with a mirror to trigger the trance state. ↩ Probably Jacques Lecoq . ↩ See page 83 of Keith Johnstone: A Critical Biography . ↩ I suppose that’s redundant. ↩ On page 51 of Keith Johnstone: A Critical Biography (it’s called “critical” but it was clearly written with Johnstone’s involvement and support, and does not seriously criticize him at any point). In The Knack , Tom gives a monologue about how to teach children to play the piano that could be lifted straight out of Impro . ↩ In 1983 Johnstone “read the riot act” to the improv players who were planning to unionize, threatening that they’d be cut out of the group for good. To quote Dennis Cahill, a group member at the time who opposed the union: “I just didn’t see the point to it. … I didn’t really see a need to confront Keith or cause Keith problems or to upset him in any way over something as simple as Who Has The Power or Who Doesn’t.” ↩

0 views
Farid Zakaria Yesterday

How to piss off your Nix friends

Warning If you are pissed off reading this blog post, I guess, mission accomplished ? Try not to take life too seriously. Unfortunately, there’s a lot worse things in life than my opinions on Nix. Seems like it’s all too easy to get people in Nix flustered, angry and out with their pitchforks. All it takes is someone proposing a markdown file for the community to lose its mind. Having been a member, exposed to and part of the Nix/NixOS community for many years, I thought I would share some personal opinions , some which are self-evident and others that are purely philosophical. Nix is brilliant and deeply flawed. Both things are true at the same time. The documentation is infamously terrible (although getting better!), the language is foreign to read for many and there are still growing pains from the governance changes. You should be free to say all of this and yet still believe Nix to be the best idea in the industry, at least this decade. The emperor has no clothes. “If we just fix the documentation and the onboarding, then there’s no stopping Nix & NixOS” - Most Nix users Nix is “having a moment”. Growth on any measurable metric is growing. There are endless posts about it. Part of human nature is the desire to win. As a result, there is a persistent fantasy that if we just fix the documentation, just smooth the onboarding, just ship a friendlier installer and so on, then everyone, even your grandmother , can be using NixOS. Nix is not meant for everyone. Nix is a power tool. Power tools have a learning curve and can occasionally take a finger. I continue to use a non-NixOS Linux machine in addition and guess what, it works well enough. It is surprisingly stable and despite what Nix leads us to believe, not everything is on fire. Our obsession with mass adoption has warped priorities and diluted the amazing possibilities Nix could allow by having to make it more palatable for broader appeal. Nix originated in Europe. It should be no surprise that it skews heavily European. We can easily affirm this from the 2025 community survey . Europeans are different from Americans. We hold different cultural values and priorities. Both have traits that I wish the other emulated. Unfortunately, where they clash however is often a point of contention. Americans are capitalist maximalists. The idea of the “American dream” is tied to it. We also have the largest military budget in the world. You’re rarely more than a degree or two from a through-line between a business and the military. This clashes with much of the European worldview. As a result, there’s a bit of an undercurrent where being an American or an American corporation is problematic in the community. Your position is suspect from the start, assumed to carry ulterior motives and it curdles into a purity test. I am clearly in the “AI is useful” camp. I have written before about how much LLMs have unlocked for me personally. The Nix ecosystem is probably the best poised for the AI-wave. I have found a newfound joy and love for my NixOS machine now with LLMs. All those weird quirks that bugged me I have been able to resolve, and declaratively reproduce for future generations. One-off AI written tools can be written and stored in my NixOS configuration with a sense of assurance that they will not collide or interfere with the rest of the system. Unfortunately there’s a loud contingent that treats AI output at best as technically unsound and at worst as some moral failing of the user. Despite nixpkgs offering AI tools, an AGENTS.md file was seen as heresy. I particularly enjoyed Linus Torvalds, on a kernel mailing list articulating better than me Linux’s position on AI: AI is a tool, just like other tools we use. And it’s clearly a useful one. […] Anybody who doubts that clearly hasn’t actually used it. These are tools we can use to push Nix & nixpkgs further. The democratic process is great for society but a software project originated from someone . Someone had the vision, birthed the idea, worked tirelessly on it and then attracted others to contribute to their ongoing vision. In the case of Nix, Eelco created Nix in 2003 as part of his PhD research. The NixOS foundation didn’t exist until 2015. That is over a decade of work towards a project driven by his own vision as the Benevolent Dictator for Life (BDFL). A non-democratic model works especially well in open-source because you are free to fork the software and try your hand at your own ideas if you disagree – the same cannot be said for our shared geography. A clear vision, whether you agree with it or not, is refreshing. Someone who can say yes or no and is not simply stonewalled by design-by-committee. DHH had said it poignantly well: “Using open source software does not entitle you to a vote on the direction of the project.” [ cite ]. Every hour spent making Nix pleasant on macOS is an hour not spent making Nix extraordinary on Linux. You would never catch an iOS developer working on Linux and yet it pains me to see those who target the Linux platform working on a Mac. For Nix, Darwin support is a bottomless tax. Closed toolchains, an SDK that shifts under you every release, a sandbox that fights you. All to chase an OS whose entire philosophy (opaque, proprietary, convention over purity) is the antithesis of Nix. When we have to target solutions that cover wildly different platforms, the end result is muddled and limited. Beauty, elegance and innovation emerge when you apply constraints and restrict a problem. Flakes are here to stay and yet their adoption is constantly brought up in order to validate its existence. I default to using it for my new projects, mostly because it’s easier at this point and it seems annoyingly tied to the new CLI format. Upon reflection though, I don’t feel like I have gained anything really over or . Subjectively my evaluations feel slower as now I’m fetching many more or jump through hoops to make every flake follow each other defeating the whole purpose of separate trees. Unless you are sharing a laptop with your family or are using a mainframe from 1980, you won’t have more than a single user on your machine. Despite this, the default installation we guide users towards is one designed for multiple users. The multiple-user install adds unnecessary complexity many don’t need or won’t understand: a build daemon, pool of users, systemd service, etc… In contrast, the single-user install is radically simpler to run, operate and triage. I don’t have to remember whether I am setting configuration for my “client” or the “daemon” and that there is a difference. Multi-user is the right default for a shared build farm. It is overkill for the single human it’s actually installed on the majority of the time. Nix is the closest thing we have to a solution where we can rebuild the entire world reproducibly. It is used by the Software Heritage Foundation as a way to reliably collect, preserve, and share all software that is publicly available in source code form [ cite ]. Some of the most amazing technology that exists in this world exists when one can make changes at multiple layers throughout the stack. This is the secret sauce to many of the hyperscalers of today. This is table-stakes for NixOS. You can implement a solution that requires new: application, compiler, library, runtime and kernel all within a single commit . Despite the ability to wildly diverge from traditional distributions, innovate and differentiate, we largely replicate the status-quo – albeit more reproducible. Whether it’s constraints imposed by needing to accommodate alternate platforms (e.g. macOS) or fearing alienating more novice-users we limit the potential of what Nix could do. Are you pissed off? Let us still be friends.

0 views

Desiring, Thinking, Knowing, Doing

This article has been posted to my Substack also. Consider these situations: You want to write a novel Someone asks you what is the capital of Norway Someone asks you how you feel today Someone is telling you that their program has a bug Suppose that you would want to write a classic program (GOFAI-style), as an agent able to implement (perform) these four scenarios. How would it work? The first scenario should implement a kind of “desire state”. That desire state should have some “target”, which could be something like: a world in which a beautiful new novel exists, or simply, a desire to produce an interesting sequence of words, by itself. You want to write a novel Someone asks you what is the capital of Norway Someone asks you how you feel today Someone is telling you that their program has a bug

0 views
Brain Baking 2 days ago

In Search Of A New Fountain Pen

Let’s start this post off with a quote from myself back in February last year when I let you all peek into my stationary drawers: I keep telling myself I don’t need more but somehow I think I do. There; the search can already stop before it has properly begun: I don’t need more. But I think I do. Shut up little devil in my head. It’s been three years since I bought a new pen and the ache is real. Especially after seeing a few top list posts appear in my RSS feed such as The Gentleman Stationer’s 12th anniversary post: what’s in your pen case these days? . Six out of ten pens are Pilots! The Well-Appointed Desk followed up with a if I could only keep 10 fountain pens post, as did Rachel’s Reflections . And now I also want to list ten. Except that I don’t have ten decent pens! As mentioned in my daily drivers post , there are three workhouse pens that are constantly inked with a few newer ones and a few that I barely use. I guess I can make up a top 5: Okay fine, a top three then. Spot four could be taken by the newer Pelikan M400 Medium ground to a cursive italic that’s fun to write with, yet I regret not getting a M600 or even M800. The 400 body is just too small. My ideal body proportion-wise is Pilot’s CH 912 that’s just perfect. The M400’s grip section is short and thin, but the M600+ versions come with a (too) hefty price increase. I guess I can still hunt down a bigger body and attempt a nib swap. My wife gifted me a Waterman Expert III Fine when I got my PhD and while it’s an excellent writer it doesn’t really come close to the three pens listed above. And then I wanted a proper Sailor with some colour but not a Slimline variant. I bought a Wancher variant on eBay that came with a busted nib so had to send that off. I postponed that for years hence it’s not really been used yet so I can’t give it a proper place in the above list. A friend listed his own top 10 as well: Pilot Falcon, Pilot CH 912 FA, Pilot CH 912 PO, Pilot Custom 823 F, Pilot Vanishing Point F, Sailor Pro Gear Naginata Togi specialty nib, CONID Bulkfiller, Montblanc 149 Calligraphy nib, and two others. After the Pilots, the price ramps up to almost which is way above my budget for a pen. I don’t really want another 912 even though the nib variability is seductive: it only comes in one colour; black. When we visited Italy in 2016, I wanted to buy a proper Italian fountain pen. The Brain Baker in me of course first thoroughly researched all Italian brands and I concluded that Aurora makes unique in-house nibs so I wanted an Aurora—perhaps even their expensive flagship model the Aurora 88. In Milan a friendly employee helped me to an Ypsilon as the others were very expensive. Two months later, some of the cheap metallic layer covering cheaper plastic came off. I also hear quality control issues for other Italian brands so I’m still unsure even though that Leonardo Momento Zero Grande special edition Tulp Fields Ru looks great. But both the size and the price are a bit too big for my liking. Is it just me or are you just paying for more looks after passing the mark? So what else is on the wish list? Back to Pilot? Yes! Their Custom 823 is a vacuum filler and in select shops in Japan you can order a flexible #15 Falcon nib along with it: impossibly difficult to import. But you can also buy a 743 with a FA nib and swap them… One for the price of two! Why not just get a single 743 then? Because it only comes in black. Except for the US-only limited edition Verdigris colour. I can almost hear the custom till going ding ding ding. I always like looking at special nibs as my daily writers are already here, lying on my desk. Sure it’s lovely to have a bit of a pen rotation going on, but I’m not dropping hundreds just to have yet another fine nib. A Platinum music nib then, which has two slits and has wide down strokes but thin side strokes? As a Belgian fountain pen writer, I’ve always wanted a CONID, especially after testing one from a friend . These expertly machined pens in small batches are made from very durable materials which… drumroll… makes them ridiculously expensive and difficult to get. But they are the ultimate demonstrator vacuum pens with swappable Bock nibs that make it easy to order a few crazy ones from say . Maybe that’s what’s missing in my lineup? The cheapest alternative are TWSBI pens, but my Eco already has plastic nicked off here and there so I’m not a huge fan of their usage of cheaper material. Perhaps the more robust Vac700R is worth taking a look at. In-between CONID and TWSBI, you have the Irish Gravitas pens, of which the new Vac 2.0 looks very attractive both price-wise and aesthetically. Vac-filled, two separate ink chambers, can be fully disassembled, and easy to screw in other Jowo #6 nibs. So what’s the conclusion here? Perhaps I already own anything I will ever need. Otherwise: I left out Visconti’s Homo Sapiens, Aurora’s 88, and Sailor’s Naginata Togi. I just can’t justify putting in so unless someone wants to get rid of theirs at a reasonable price, those will stay within the bounds of distant dreams. Whatever you do, don’t look at the latest Sailor and Pilot prices. Pilot is planning to increase their prices in October—again. On top of that, their limited edition CH 912 is priced at which is more than twice the base price. And Sailor? They’ve been upping the stakes and marketing themselves as more premium: since 2024, prices have literally doubled. Yet another reason to hold your horses… Or to get out that cheap but very serviceable Kaweco Sport? Related topics: / fountain pen / By Wouter Groeneveld on 18 July 2026.  Reply via email . Pilot Custom Heritage 912 Waverly Lamy 2000 Fine Pilot Vanishing Point Matte Black Fine Buy a too expensive and too big Leonardo just for the looks Import a too expensive Pilot 743 with FA nib in Verdigris Buy a Gravitas Vac 2.0 and order another Jowo semi-flex nib to fiddle with Meanwhile be on the lookout for a Pelikan M600 body? If all else fails, get a Platinum 3776 with a music nib?

0 views
Jim Nielsen 2 days ago

Make It Work vs. Make It Good

There are two wolves inside of me, lol. Some days I want to be a “designer”. Other days I want to be a “developer”. On the days I find myself wanting to feed the developer, it’s often because making something “work” seems easier (and more impressive) than making something “good”. Making something function often results in a reaction of “Wow, that’s so cool! It didn’t work before and now it does! And I could’ve never made that, nice job!” And sometimes it’s like, good job, you made a bear ride a unicycle . Not really what bears are supposed to do — and they’ll probably never be good at it — but it’s novel and functioning! However, the task of making something good — of arriving at a solution that is obvious — is often met with a kind of ambivalence, like “Nice work…I guess? Seems obvious tbh.” That’s the work of design: to make something so good, it’s obvious. But there’s often little acclaim for the obvious because, well, it’s so obvious (in hindsight). This plays out in many different ways. For example, consider a task like making a web site responsive. In my experience, it’s often quite easy to get people to say “Hey that’s cool, it looks like a mobile site now! Good job!” Getting to that point is often just a matter of sticking a few media queries in your CSS. And people are impressed because they’re not honing in on the details of how it works, just that it works at all. “Cool, the site displays on a mobile phone now! We can move on.” But just because it works doesn’t mean it’s good. And that extra mile to “it works on mobile and it’s also a good experience” is a ton of work. Is it fast? Is it accessible? Is it intuitive? Does it work across multiple devices? Can it be iterated on quickly? So. Many. Questions. “Does it work?” is a binary question. “Is it good?” is a subjective question whose answer lives at the intersection of multi-disciplinary knowledge and taste, which is to say: it’s harder to answer than “Does it work?” “Let’s do X” often boils down to two stages: To “make it work”, all you gotta do is get it running. Consensus on when to applaud and reward the work is simple because it’s either working or it’s not. To “make it good” requires all kinds of nuanced work. Consensus on when to applaud and reward this work is often impossible to discern because not everyone agrees on what “good” looks like. “Make it work” is the first 90% of the work. “Make it good” is the other 90%. Reply via: Email · Mastodon · Bluesky Make it work Make it good

0 views
Max Woolf 2 days ago

What's the deal with all the random weekly quota resets for agents lately?

Subscription-based coding agents such as Claude Code and Codex famously have 5-hour and weekly quotas on their LLM usage. Both of these are understandable: 5-hour quotas help stagger usage so the servers don’t get overloaded, and weekly resets prevent users from dumping an entire month’s worth of usage into a single day which a) also prevents overload and b) stops the user from just unsubscribing after they do so. Both Anthropic and OpenAI have played around with quota limits, from doubling them for a limited time to even removing the 5-hour quota. These model providers can reset the weekly quota for all users, often gifted as compensation in the event of technical glitches on their end. If, for example, you have a $100/mo Codex plan, then a weekly reset is worth $25 to you assuming you fully consume your quota—nowadays with the cost of top-tier LLMs like Fable 5 and GPT-5.6 Sol , that’s easier to do. However, these quota resets are not telegraphed and are generally not announced through company-owned channels: unless you follow specific people such as Thibault Sottiaux (Tibo) , the engineering lead for Codex, you just look at your weekly quota and see it’s at 100%. You can even get a quota reset hours before your actual reset and not benefit from it all. Recently, in the wake of the release of Fable 5/GPT-5.6, Anthropic and OpenAI have been doing weekly quota resets for their harnesses far more frequently. In the past two weeks, OpenAI has directly reset the Codex weekly quota six times : July 9 , July 10 , July 10 (again) , July 14 , July 15 , and July 17 . That’s not even getting into the rarely discussed banked reset system for Codex, where OpenAI gave quota resets July 12 and July 13 which can be manually used at any time but expire within 30 days. https://codex-resets.com tracks Codex resets. No one wants to be the weirdo who complains about literally getting free stuff because if the quota resets stop, they will be the one blamed for it. Despite that, I’m going to be the weirdo who complains about literally getting free stuff. As a person who doesn’t like wasting money if I can easily avoid it, I try to use as much of my weekly quota as I can. I was on the $20/mo Codex plan but with the promise of GPT-5.6 on the horizon and with the frequency I kept hitting the 5-hour limits, I upgraded to the $100/mo plan. By using GPT-5.5, I used the 5x prompt capacity to build and test a number of ambitious projects, but that’s a topic for another blog post. I was able to consistently exhaust my weekly quota over the full week, although I had to have a browser window open with my Codex usage to constantly monitor it. I’ve set reminders on my phone for the exact time a 5-hour or weekly quota resets so I can keep running more prompts—as an aside, I wish there was a canonical platform endpoint to programmatically check Codex usage amount and the quota reset times so I could just vibecode an app to manage for me. Before the release of GPT-5.6, I thought “should I deliberately exhaust my quota all on the weaker GPT-5.5 and gamble that OpenAI does a quota reset to greet GPT-5.6?” (I did not exhaust my quota and OpenAI did indeed do a quota reset) GPT-5.6 Sol is indeed a great model that does live up to the hype, and I’ve had to create new projects just to have an excuse to test its limits. The rate of quota usage is about the same as GPT-5.5 even with a prompt running at all times. Despite that, almost every time a reset has happened after the GPT-5.6 release, it has been when my weekly quota has been at 50+%, which makes me feel like I “wasted” $12. There’s also the factor that when the quota resets occur, the weekly reset time is unset, so I have to input some prompt just to trigger it again even if I’m doing something else. @emollick / X After the flurry of quota resets over the past two weeks, the intended excitement has instead become annoyance as I now have to urgently create new ideas to spin down the quota before it inevitably resets again . Random rewards are supposed to give a dopamine hit but I end up with a net dopamine deficit from both the sense of wasted quota and having to abruptly change my plans. I admit that this may just be a sign I’m burnt out and need to take a break. Tibo polls if there are too many resets , with a few thousand responses. Resets of the weekly quota for all users must be ludicrously expensive for these companies, although when you already spend billions of dollars in CapEx per year it’s likely a rounding error. The recent surge of quota resets is likely not a coincidence: July has been an insane month in LLM releases, with not just Fable 5 and GPT-5.6 Sol pushing frontier models even further, but also Grok 4.5 , Muse Spark 1.1 , and Kimi K3 offering more options across the cost/utility curve. The cynical take is that weekly quota resets are not intended to be fun serendipity, but instead intended to prevent power users from experimenting with sufficiently competitive competitors once the quota naturally runs out. I don’t expect weekly quota resets to last forever even if competition intensifies, because if quotas keep resetting this frequently they won’t matter at all. It would instead give me an incentive to downgrade from the $100/mo plan back to the $20/mo plan to avoid wasting quota, which I don’t think is OpenAI’s intended goal.

0 views
Unsung 2 days ago

The Swiss Cheese model, pt. 1

Have you head of the Swiss Cheese model ? You see it sometimes in descriptions of how complex systems fail. The visual usually goes like this: = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/the-swiss-cheese-model-pt-1/1.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/the-swiss-cheese-model-pt-1/1.1600w.avif" type="image/avif"> The whole idea is: even if you have multiple layers of safety – like many slices of cheese – there are always holes in each slice. Typically, a hole in any slice is covered by a non-hole in the previous one or the next one. (For example, a car might not allow you to grab your keys if you have not shifted to park – or, if you start driving with a handbrake on by accident, the car might yell at you.) But occasionally, the holes just happen to line up, and a larger disaster strikes. The model is used in analyses of past accidents, and prevention of future ones. It has proponents and detractors. It’s hard to talk about its applications because the most common examples are horrific. In my book, I wrote about Therac-25 , and that was a really unpleasant chapter to research and to write. Other go-to case studies are equally bleak: Chernobyl , Challenger , the Tenerife airport disaster , the Deepwater Horizon explosion . But I wanted to share it because in my head it applies to UI design also, and sometimes helps me think of how small details add up to a larger whole. In this first part, let’s start with a more traditional example, although a non-drastic one. Here’s a story of Knight Capital Group, a financial services and trading firm. I’m going to hand off the summary of the accident to Henrico Dolfing : On the morning of August 1, 2012, Knight Capital Group opened its systems for what should have been a routine trading day, yet within minutes the firm began sending a flood of unintended orders into the U.S. equity market, buying high and selling low across dozens of stocks in a pattern that made no economic sense and could not be stopped through normal controls. What initially appeared as unusual market activity quickly escalated into a systemic failure inside one of the largest market makers in the United States, with algorithms behaving in ways that neither traders nor engineers could fully understand in real time. […] By the time the issue was identified and the system shut down roughly 45 minutes later, Knight had generated more than 4 million executions across 154 stocks, covering approximately 397 million shares, and accumulated positions worth billions of dollars, resulting in losses of more than $460 million […]. The scale of the incident was not only financial but structural, as a single deployment failure had propagated through a system responsible for a meaningful share of U.S. equity trading. […] Here’s what happened. Long ago, the firm built a pretty boring function called Power Peg to automate some transactions. The function used a standard shared limiter that made it stop executing when all the required transactions were fulfilled. After some years in use, the function was deprecated in 2003 and stopped being used then, but crucially, the code was never actually removed. Some time in between 2003 and 2012, the limiter functionality was upgraded, and the old code stopped being compatible with it. All code in production was rewritten to use the new limiter, but the Power Peg feature wasn’t, as it was already deprecated and not in use. In 2012, the firm started writing a new program for automated transactions. Its creators decided to reuse the same software flag that previously activated Power Peg. The existence of the old code was known at the time, and the idea was that the code would be overwritten by the new program, so the reused flag would only trigger new code. In July 2012, the new code was finished and the firm started to install it on all the servers, overwriting Power Peg. The firm intended to deploy the new code to all eight servers, but a mistake resulted in it only arriving on seven servers. No one caught that mistake. At this point, you can piece it all together. At this point, I imagine many of you have been wincing more and more with each passing paragraph. On August 1, the flag for the new functionality was turned on. Everything was fine on seven servers, but on the eighth one, the flag reawakened dormant code that immediately started executing. As the old code was not compatible with the new limiter, it was never limited, cascading into millions of transactions in less than an hour, and a lot of collateral damage; the subsequent market reaction to the news of the firm losing over $400 million and caused its own stock to tank. Oh yeah, I didn’t mention this yet – the failure killed the firm. (Well, it resulted in a merger, but that seemed to be a way to save face after Knight Capital Group almost went under.) How does the Swiss Cheese apply to this? You can see it as five holes in five different slices of cheese: What’s important to understand about this model is that either of these in isolation would objectively be a small mistake, and caught by the other slides. As a matter of fact, any four of these happening would still not add up to a catastrophe. But in this case, the five mistakes lined up perfectly. Many analyses of such accidents blame a single event in the chain – in this case often the sysadmin that didn’t deploy the code to eight servers properly – but this is primarily because we like stories of individual agency, and are not well equipped to understand stories of systems . (Even Star Trek added a Borg Queen, after all.) That’s the accident eventually dubbed “a Knightmare.” More examples of systems closer to our hearts in following parts. #bugs #definitions it was a mistake to not actually remove the old code it was a mistake to reuse the same flag it was a mistake to not deploy to all 8 servers it was a mistake to not have a procedure for someone to double check the deployment it was a mistake to not have a way of auto-detecting (and perhaps auto-stopping) the runaway processes when the regular limiter failed

0 views
Ahead of AI 2 days ago

Controlling Reasoning Effort in LLMs

It has been almost two years since OpenAI released o1, a model that popularized the idea of LLM-based reasoning models. DeepSeek-R1 followed about four months later, together with details of a reinforcement learning with verifiable rewards (RLVR) recipe to train such reasoning models. Last week, OpenAI released the GPT-5.6 model family. It comes in three sizes, each with roughly five or six reasoning-effort settings. Figure 1: The GPT 5.6 Sol model with different reasoning effort settings. (Benchmark numbers for Ultra are currently not available but should be relatively similar to Max, since it uses a similar effort level but accelerates the work with four subagents.) So yes, reasoning models are here to stay. They have become a standard part of modern model releases. In the past, I covered the methodology of reasoning models ( Understanding Reasoning LLMs ) as well as relevant research papers ( The State of Reinforcement Learning for LLM Reasoning and The State of LLM Reasoning Model Inference ). And I even wrote a whole new 440-page book on how to develop reasoning models, Build A Reasoning Model (From Scratch) . Figure 2: My new Build A Reasoning Model (From Scratch) book. In color! These resources have focused on turning a conventional LLM into a reasoning model. Now, in this article, I want to focus on and explain how to develop a reasoning model that has multiple effort modes, similar to what’s shown in the figure at the beginning of this article. No worries, this article can be read as a standalone article. However, the aforementioned resources may be interesting and useful. When talking about pretty much any machine learning or AI technique or subfield, the one lesson is that we usually shouldn’t take technical terms “literally”. For example, an (artificial) neural network in machine learning and AI doesn’t literally work like a biological neural network like the human brain. Similarly, when talking about “reasoning models”, we shouldn’t expect that these models literally reason like us humans. In the context of AI and LLM research, “reasoning model” means a model that outputs an intermediate reasoning trace, which is like an intermediate response that works through a question or task step by step. It’s probably easiest to explain this by showing an example. Figure 3: Illustration of a conventional LLM answer (left) and an answer by a reasoning model (right). There are essentially two ways to improve (reasoning) task performance: training scaling and inference scaling. Figure 4: Training and inference-scaling are two ways to improve LLM and reasoning model problem-solving capabilities. Plot based on Learning to reason with LLMs Let’s briefly talk about training first. In a nutshell, DeepSeek-R1 proposed training an LLM using reinforcement learning with verifiable rewards (RLVR) to turn it into a reasoning model. RLVR is a technique to provide a reward signal ( and ) for verifiable data domains. These verifiable data domains here are math (we can use a symbolic math checker like SymPy or WolframAlpha to check results) and code (we can use a compiler or unit tests, or integrated platforms like LeetCode) to check for correctness. Figure 5: Illustration of accuracy and format rewards during RLVR training. Notably, the reasoning trace itself was not used for training or updating the model. Although they tried to use this intermediate response information for training, the DeepSeek-R1 paper reported that it wasn’t helpful for the model training, so it was ultimately not used. (Whether and how to incorporate intermediate reasoning traces in the training signal via process reward models is an active area of research.) Figure 6: The intermediate reasoning trace is ignored during RLVR; only the final answer and response format determine the reward. Anyway, just training on the output rewards alone, as Figure 7 shows, turned out to be sufficient for the model to learn how to reason through a problem, meaning that it would learn to write intermediate explanations, backtrack, and self-correct itself. These moments when the model realizes that it made a mistake and self-corrects itself are called “Aha” moments. Figure 7: An example of an aha moment, where a reasoning model notices an error in its intermediate reasoning and corrects it before producing the final answer. By the way, while DeepSeek-R1 is inarguably the more popular paper, and the paper that created excitement around reinforcement learning with verifiable rewards and the development of reasoning models, there is another paper, Kimi K1.5 , published on exactly the same day on arXiv (22 Jan 2025). Also, the term RLVR was already coined two months earlier in Tülu 3: Pushing Frontiers in Open Language Model Post-Training . One reason why the DeepSeek R1 is ultimately the more popular paper is that it demonstrated that reasoning behavior can be achieved with pure reinforcement learning (RL). Figure 8: DeepSeek-R1-Zero applies RLVR directly to the pretrained base model without supervised fine-tuning. For instance, Tülu 3 and Kimi K1.5 applied reinforcement learning on top of a supervised fine-tuned (SFT) model. The DeepSeek-R1 model was also trained from an SFT checkpoint of the DeepSeek-V3 base model, and it included a DeepSeek-R1-Zero variant trained with pure RLVR. R1 Zero is a weaker model than R1, but it showed that RLVR is sufficient for teaching the model to generate and use reasoning traces. ​While R1-Zero was more of a proof-of-concept model, note that the full DeepSeek-R1 reasoning model training pipeline is usually multi-stage and a bit more complicated, as mentioned above. Figure 9: More detailed reasoning model training pipeline. This one depicts the various DeepSeek-R1 models. For more details, see my other article: Understanding Reasoning LLMs ​ By the way, most of today’s LLMs are effectively reasoning models, meaning they have been trained in a similar fashion to DeepSeek-R1 using a form of RLVR. Next to improving reasoning behavior through training, another lever for improving model performance is inference compute scaling. In short, this means that we are spending more compute after training the model, during usage, to get better answers. This is a whole topic by itself, and you could read through my The State of LLM Reasoning Model Inference for a more detailed rundown: I will try to summarize what’s most essential to mention as background info below. First, training a model with RLVR is already implicitly leading to a form of inference scaling, since reasoning models usually output more tokens during inference compared to conventional LLMs, and that means we are spending more compute during inference. Second, we can further adjust this output length via reasoning effort levels, but more on that later. ​Third, there are many additional inference scaling techniques. A popular one is self-consistency, which is often implemented as a form of majority voting where the model is queried multiple times, and the final answer is selected via majority vote. Figure 10: An example of self-consistency, a popular inference scaling technique. This can be applied to conventional LLMs as well as reasoning models. Also, this method can be used on demand and in addition to reasoning training. A good example of that is DeepSeekMath-V2, where the researchers applied extreme inference-scaling on top of a reasoning model (specialized for math) to achieve state-of-the-art performance on challenging math olympiad-type problems. Figure 11: Two types of inference scaling (self-consistency and self-refinement) used together to improve math performance. Figure adapted from DeepSeekMath-V2: Towards Self-Verifiable Mathematical Reasoning But again, I will refer to my other article, The State of LLM Reasoning Model Inference for an overview of other techniques: You may have seen the tokens in the earlier “Aha moments” figure. I also included the corresponding figure below so you don’t have to scroll all the way up. Figure 12: Common formatting tokens in reasoning models. These and tags are cosmetic with respect to reasoning ability. They do not make the model reason, and they are not required to achieve good reasoning performance. One could train the same model without these delimiters and likely reach similar benchmark performance. The purpose of these tags or tokens is mainly to mark where the reasoning trace begins and ends so that the training pipeline or user interface can separate it from the final answer and optionally hide it from the user. (UIs like ChatGPT or Codex usually do this.) The point here is that the tokens are not giving the model the ability to “think” or reason or reason better. One could train the same models without such tokens and reach similar benchmark performance. There is also nothing special about the literal strings and . Another pair of delimiters could serve the same purpose. By the way, the way this is implemented is typically by adding a formatting reward during the RLVR stage. So instead of just rewarding the model based on answer correctness, one would provide additional reward for the use of <think> tokens, which in turn encourages the model to use those. In DeepSeek-R1, for example, the overall reward was calculated as where the format reward was a simple rule-based check that encouraged the model to place its reasoning inside: reasoning trace The first generation of reasoning models was dedicated reasoning models. With that, I mean that there was a DeepSeek-V3 base model and a separate DeepSeek-R1 reasoning model. No matter what the prompt is, R1 generally outputs very verbose responses using lots of tokens, even for simple prompts. It also lacks a built-in option to turn off the reasoning mode. Figure 13: Reasoning models are very verbose, even for the simplest prompts. Later models, like Qwen3 and others, experimented with hybrid approaches, where the same model can behave like a regular instruction fine-tuned model or a reasoning model on demand. Note: Some model developers call this “thinking mode,” while others call it “reasoning mode.” Both terms refer to the same behavior. In Qwen3, this is handled via the tokenizer using or . Under the hood, setting essentially adds an empty section to the beginning of the assistant response to turn off Qwen3’s reasoning (”thinking”) mode. Figure 14: Response of Qwen3 0.6B reasoning model with and . (The empty tags are hidden in the interface on the left as they are part of the modified input prompt, not the generated answer.) How is this implemented during training, such that the model supports this toggle during inference time, as shown in the figure above? In short, as explained in the Qwen3 technical report , this on/off behavior is introduced primarily through supervised fine-tuning (SFT) and then reinforced during general RL in their largest flagship models. For instance, after the initial reasoning model is trained via long-chain-of-thought SFT and reasoning RL, they add a “Thinking Mode Fusion” stage. During this additional SFT stage, the model sees both thinking and non-thinking examples: Thinking is the default behavior, so /think can also be omitted. The subsequent general RL stage further reinforces this mode and format following. These and flags are a “soft” switch. However, the setting mentioned earlier, which force-adds the empty in the False case, acts then as a “hard” switch. Figure 15: “Thinking Mode Fusion” in Qwen3’s training pipeline to enable the reasoning mode on and off switch. In other words, the tokenizer does not add to the query. It directly fills in the empty section at the beginning of the assistant response. The model only sees the resulting tokens and continues directly with the answer. Anyway, this on-and-off toggle is essentially a simplified version of the reasoning effort levels in GPT-5.6 and others, which I cover in the next section. In this section, I want to provide a brief overview of how the different reasoning effort toggles may be implemented, which have been introduced in models like GPT 5 and are present in pretty much any flagship model today. Concretely, at the beginning of this article, I showed a figure from the Codex GPT 5.6 interface that lets users select multiple reasoning “effort” settings. Figure 16: GPT-5.6 exposes six reasoning effort settings, ranging from Light to Ultra. The following subsection will illustrate how these settings may be implemented. Then, in the next section, I will go over some of the more interesting research papers related to this topic. Unfortunately, the implementation details of their effort settings are not shared by OpenAI, but there is some evidence out there that can be used for educated guesses. For instance, via their open-source gpt-oss models from last year (I wrote about them in From GPT-2 to gpt-oss: Analyzing the Architectural Advances ), we know that OpenAI allows us to toggle the reasoning effort setting via the system prompt (”Reasoning effort: low/medium/high”) that is prepended to each prompt. Figure 17: The gpt-oss chat template inserts the selected reasoning effort into the system message before sending the prompt to the same model. As expected, the reasoning effort directly affects the response length and accuracy, as shown below. Figure 18: Response length and quality of gpt-oss models under different reasoning efforts (annotated figure from the model card ) Presumably, their GPT 5 models, including the recent GPT 5.6 models, use a similar approach. By the way, note how different effort settings scale the response length in the figure above. The effort level seems directly correlated to token usage, which in turn seems correlated to accuracy. It might be possible to come up with effort settings beyond the “high” one, but I assume performance would saturate at some point. This saturation can be seen more clearly for the GPT 5.6 Sol model, which also shows that increasing reasoning budgets can become uneconomical at some point. Figure 19: Reasoning effort increases both API cost and coding-agent performance, with diminishing returns at the highest GPT-5.6 settings. Figure based on the Artificial Analysis Coding Agent Index v1.1. Another good, very recent data point that shows the relationship between reasoning effort, token usage, and benchmark performance is this week’s new open-weight Inkling release by Thinking Machine Labs. Figure 20: Increasing the Inkling effort level generally increases generated tokens and benchmark performance, with diminishing or uneven gains at higher effort. Figure from the Inkling announcement blog . As discussed in this section, during inference, the reasoning effort level can simply be controlled via a system prompt. (The ChatGPT UI presumably simply maps the menu choice to a system prompt.) However, this would not work for an arbitrary model and requires certain modifications to the training pipeline, which will be discussed next. While the training details are not public, neither for GPT 5.6 nor the open-source gpt-oss models, typically, the reasoning effort label is included in prompts during post-training. There are typically two ways to implement this. First, we can implement it as part of the RLVR process and apply a different length penalty when different system prompts are used. For example, a high length penalty when “Reasoning effort: low” and a mild or no penalty when “Reasoning effort: high”. Second, we can fine-tune the model after RLVR to follow different effort instructions via supervised fine-tuning (SFT). For instance, after the core RLVR stage and during SFT, the prompts in the training dataset are paired with target responses that exhibit the desired amount of reasoning. (The targets may be written by humans, generated by another model, or generated and then filtered.) Figure 21: Illustration of effort-conditioned RLVR and SFT. (This is a possible implementation, not a confirmed description of OpenAI’s training pipeline.) During this SFT stage, the model learns the association between the effort label and the target reasoning length directly from the training examples. An RL-based implementation would instead place the effort labels and budget-aware reward inside the RLVR stage. The two approaches could also be combined, which I suspect was done for both gpt-oss and GPT 5.6 (note that effort settings in GPT 5.6 are likely just changing the system prompt for a given user query). The just-released Inkling technical report gives a small but somewhat concrete example of effort-level training. Figure 22: Inkling sweeps a continuous effort value between 0.2 and 0.99; higher effort generally produces longer responses and higher benchmark scores. During large-scale RL, they did two things for each sample: Specified the desired effort level in the system message. Adjusted the cost assigned to each generated token. Conceptually, the reward likely looked something like this: Here, e is the requested effort level and λ(e) controls the token penalty. Low effort uses a larger per-token cost, encouraging shorter reasoning traces. High effort uses a smaller per-token cost, allowing the model to spend more tokens. Then, at inference time, Inkling receives a system message such as Thinking effort level: 0.8, and adjusts its token usage accordingly. The difference between Inkling and models such as gpt-oss and GPT-5.6 is that the effort label is a continuous number between 0 and 1 instead of ordinal labels such as low, medium, and high. This places Inkling’s effort-level conditioning primarily in the Reasoning RL stage, not only in the later SFT stage. They do not disclose the exact reward formula, token-cost coefficients, or whether effort conditioning was also included in SFT, though. Before moving on to the reasoning effort papers, I want to briefly connect this section back to the earlier “2.3 Inference scaling in a nutshell” section. Earlier, I separated scaling into training compute scaling and inference-time scaling. The GPT-5.6 interface provides a nice way to illustrate the difference, as shown below. On the left, selecting Luna, Terra, or Sol changes the model itself. As a rough analogy, this corresponds to training compute scaling. These are separate trained models. At a fixed training recipe and dataset size, a larger model requires more training compute. It also generally requires more compute per generated token. On the right, we keep the model fixed and only change the reasoning effort. This is inference-time scaling. The model weights stay the same, but the model is allowed to spend fewer or more tokens working on the answer. Figure 23: The model selection and reasoning effort menus correspond to two different scaling axes. Selecting Luna, Terra, or Sol changes the model, whereas changing the reasoning effort adjusts the inference-time compute for a fixed model. One small terminology caveat is that selecting a different model from the menu is not training scaling at that moment. The training has already happened. It is better to think of the model menu as selecting among models that were produced at different training scales. The Artificial Analysis results below show how these two axes interact in practice. Each blue curve corresponds to one model, Luna, Terra, or Sol. Moving along a curve by increasing the reasoning effort is inference scaling. Moving from one model curve to another corresponds to model scaling, which I use here as a practical proxy for training scaling. As expected, both approaches can improve the benchmark score, but they also increase the cost. More interestingly, the curves overlap. For instance, a smaller model at a higher reasoning effort can sometimes reach a similar score as a larger model at a lower reasoning effort. Figure 24: Training scaling and inference scaling for the GPT-5.6 model family on the Artificial Analysis Coding Agent Index. Moving along each model curve corresponds to increasing the reasoning effort. Moving across the Luna, Terra, and Sol curves corresponds to selecting a different model. By the way, the x-axis in this figure shows API cost rather than raw compute. The API cost is a useful practical measure, but it also depends on the provider’s pricing and the number of generated tokens. Also, the exact shape of these curves is benchmark-specific. So, the model size and reasoning effort form two separate knobs. We can use a larger model, increase the reasoning effort, or combine both. Which combination is best depends on the desired accuracy, cost, and latency. So far, the article should give you a pretty solid understanding of how reasoning effort modes work and how they are implemented. This is a fine point to wrap the article if you are short on time. Otherwise, if you want to look into some of the nitty-gritty details of some of the recent open-weight models, please read on! [This section is fine to skip unless you are interested in some additional details] Section 5 described two possible ways to train reasoning-effort controls, namely effort-conditioned supervised fine-tuning and reinforcement learning with different token costs. Originally, I wanted to cover research articles on alternative ways to implement reasoning budgets. However, reading through most of these articles, they seemed more like proofs-of-concept that may or may not work well in practice. So, instead of covering those, I decided to pivot a bit and cover those recipes used by state-of-the-art and notable open-weight (flagship) LLMs. For these models, there is at least evidence that the methods work in practice. This leaves six examples. DeepSeek V4, Nemotron 3 Ultra, Kimi K2.5, GLM-5, Qwen3, and Inkling. They have different levels of detail in their reporting, but each contributes a useful variation. (I exclude models whose reports only show an effort setting in the user interface without explaining how that behavior was trained.) Let’s start with the DeepSeek V4 technical report , which describes the use of three modes: Non-think produces a direct response without a reasoning trace. Think High is the classic approach where the model places the reasoning trace between <think> and </think> tags. This is similar to what was discussed in the DeepSeek R1 section (section 2) at the beginning of this article. Think Max is the same as above but adds a special system instruction. (More on that below.) The additional system prompt instruction for Think Max starts with “Reasoning Effort: Absolute maximum with no shortcuts permitted.” Figure 25: Reasoning effort control overview from the DeepSeek V4 documentation At first, this sounds like a simple prompt engineering trick, but this prompt is actually backed by a different training setup. That is, each mode uses its own context window and length penalty (unfortunately, the exact length penalty implementation is not detailed in this report). Think Max receives a longer context window and a smaller length penalty than Think High, which gives it more room to continue reasoning. So, the system instruction selects a behavior that was created during post-training. Adding the same instruction to an arbitrary model would not have the same effect. Figure 26: DeepSeek V4 describes the three effort modes and the larger teacher pool in separate parts of the report. The teacher pool contains more than ten domain specialists. The report does not disclose how these teachers map to Non-think, Think High, and Think Max. Unfortunately, the public, and otherwise very detailed DeepSeek V4 report does not connect the descriptions of the reasoning mode and domain specialists in enough detail to reconstruct the exact teacher assignment. However, the report states that the final model, which supports different reasoning effort levels, was created via on-policy distillation from said teachers. To summarize, DeepSeek V4 develops the three reasoning specialists during post-training. Starting from the base model, it applies supervised fine-tuning followed by RLVR via GRPO. The RL configuration differs for each mode. In particular, each specialist uses its own context window and length penalty, while Think Max additionally receives a special system instruction. Then, including domain specialists, the different reasoning mode specialists are distilled into a single checkpoint that supports all three effort modes. The Nemotron 3 Ultra technical report describes three settings called reasoning-off, regular, and medium-effort, analogous to DeepSeek V4 in the previous section. Medium-effort is the cheaper reasoning mode compared with regular. NVIDIA introduces this mode during SFT using examples generated by GPT-OSS-120B in its medium-effort mode, and then further optimizes it during RLVR. About 2.5% of the RLVR prompts use medium-effort (this corresponds to length-based adjustments applied to their rewards). At inference time, all three modes are selected through the chat template . Figure 27: Nemotron 3 Ultra reasoning settings via the chat template (examples from the official model card ) 1) Regular is the default and uses , which starts the assistant response with an opening tag. 2) Medium-effort uses together with where the latter setting also appends {reasoning effort: efficient} to the latest user message. By the way, to further complicate things, the regular and medium-effort modes can also be combined with a separate inference-time reasoning budget. This budget acts as an external stopping mechanism. In the released implementation, the chat client asks the model to end the reasoning trace near the chosen token limit. If the model has not emitted , the client closes the reasoning block and continues generation to produce the final answer. The learned effort mode determines how the model uses its reasoning tokens, while the budget constrains how long the reasoning trace can continue. This makes it possible to pair either mode with a tighter or looser budget depending on the desired cost and accuracy. 3) Reasoning-off uses , which prefills an empty block (similar to Qwen3 discussed in section 4) so that the model proceeds directly to the final response. Thus, these are chat-template controls rather than system prompts. The inference controls described above are backed by two related SFT components. The first introduces medium-effort behavior using GPT-OSS-120B traces, as discussed earlier. The second prepares the model for hard reasoning budgets. To construct this training data, the authors take regular reasoning traces, truncate them at randomly selected token budgets, and keep the original final answers. The inserted token is masked from the SFT loss. As a result, the model sees examples where it has to move from an incomplete reasoning trace to the answer after the reasoning block has been closed externally. Medium-effort training then continues during RLVR. About 2.5% of the RL prompts use the medium-effort setting across math, STEM, and coding tasks. The report notes that the mode can be calibrated through reward hyperparameters where length-based reward adjustments provide additional control over the cost-quality trade-off. Figure 28: Nemotron 3 Ultra introduces medium effort with teacher-generated SFT data, random-budget truncation, and a small medium-effort subset during RLVR. The Kimi K2.5 technical report discusses a training method called Token Efficient RL for lower reasoning effort. (While there was a K3 announcement this week, the reasoning-effort methodology of K3 is not publicly disclosed, but it could be similar or related to K2.5.) The report mentions that a fixed token budget can make a reasoning model overfit to short solutions. That means the model becomes more concise (i.e., faster and cheaper), but it may lose the ability to benefit from additional inference-time compute and can thus perform poorly. Figure 29: The proposed Toggle method makes Kimi K2.5 much more token-efficient while keeping the overall benchmark performance similar. Annotated figure from https://arxiv.org/abs/2602.02276 Kimi K2.5’s method, called Toggle, alternates between two RL phases every fixed number of training iterations: 1. In the budgeted phase, correct solutions are encouraged to stay within a problem-specific token budget. 2. In the unconstrained phase, the usual maximum generation length is restored so that the model can still learn from longer solutions. For each problem, the budget is estimated from a selected percentile of response lengths among correct rollouts in RLVR. The budget constraint is then only activated once the mean accuracy on that problem exceeds a threshold. This avoids forcing the model to shorten its reasoning before it can solve the problem reliably. Figure 30: Overview of the two phases of the Toggle method. The report evaluates Toggle on K2 Thinking and finds that it reduces generated tokens by about 25 to 30% with little change in benchmark performance. The same behavior also transfers from math and coding RL tasks to GPQA and MMLU-Pro. Toggle supplies a concrete flagship-model recipe for training a more token-efficient reasoning policy while preserving its ability to scale at test time. Toggle operates entirely during RL training. Both alternating phases update the same policy (i.e., LLM), and the final (unified) checkpoint has no budgeted-versus-unconstrained selector. At inference, the resulting model then runs in thinking mode by default. Interestingly, though, Kimi K2.5 itself exposes a separate binary choice between thinking and instant modes in some APIs I checked (like vLLM or SGLang). Thinking mode is enabled by default. Instant mode disables the reasoning trace through thinking: in the official API or when serving the model through vLLM or SGLang. However, these settings are separate from Toggle. Also, the official Kimi report does not provide a separate training recipe for instant mode. However, K2.5’s SFT data were generated using both the earlier K2 model, which produces direct responses without long reasoning, and K2 Thinking, which produces extended reasoning traces. This likely exposes the unified checkpoint to both response formats similar to what’s done in Nemotron 3 above. At inference time, the chat template selects between them by prefilling either an open <think> tag for thinking mode or an empty block for instant mode. But again, unfortunately, the report does not disclose the exact data mixture or whether additional mode-specific RL was used. The newer Kimi K3 provides a more direct inference-time effort interface. The current Kimi Code documentation lists three settings called low, high, and max, with max as the default. These are passed through the parameter. However, Moonshot has not yet explained how the three effort levels were created during training. Its launch post says that these details will appear in a future K3 technical report, so I’ll stay tuned for that. The GLM-5 technical report extends the binary on/off thinking switch introduced with GLM-4.5 to multi-turn and tool-using scenarios. It describes three related behaviors (rather than three effort levels): Interleaved thinking: this inserts a reasoning block before each response and tool call. Preserved thinking: here, the chat retains earlier reasoning blocks across turns so that the model can reuse them later. Turn-level thinking: this enables or disables reasoning separately for each request in a conversation. At inference time, turn-level thinking is the actual on-off switch. In the Z.ai API , thinking is enabled by default and can be disabled for an individual request with thinking: . The hosted implementation is not disclosed but the open GLM-5 chat template shows the equivalent mechanism when self-hosting with Transformers, vLLM, or SGLang. It starts the assistant response with when thinking is enabled and when it is disabled. The latter closes the reasoning block immediately, so generation proceeds directly to the final answer. The report says that these behaviors are introduced during multi-task SFT together with an updated chat template. After SFT, GLM-5 goes through reasoning RL, agentic RL, and general RL. And a final on-policy distillation step uses checkpoints from the preceding stages as teachers. This helps the final model recover capabilities that may have weakened during the sequential RL stages. Figure 31: GLM-5 training pipeline. Qwen3 was already covered in Section 4, so I will only summarize the parts that matter for this comparison. According to the Qwen3 technical report , its post-training pipeline has four stages. These are long-chain-of-thought SFT, reasoning RL, Thinking Mode Fusion, and general RL. Thinking Mode Fusion is the key stage for the effort on-off switch. Here the model is trained via SFT on a mixture of thinking and non-thinking examples. The /think examples contain a reasoning trace, while examples begin with an empty block that is accompanied by a short answer. The following general RL stage reinforces instruction and format following for both behaviors. Qwen3 also supports a hard thinking budget. At the requested threshold, the reasoning span is stopped and a stop-thinking instruction is inserted before the model continues with its final answer. The report says that this partial-reasoning behavior was not trained explicitly. It emerged after Thinking Mode Fusion. This gives Qwen3 a learned on-off switch plus an inference-time budget. It is similar but simpler than the DeepSeek V4 and Nemotron recipes. Inkling was already discussed in Section 5.3. The short version is that its technical report mentions that they use continuous effort conditioning (values between 0.0 and 1.0) rather than fixed effort labels. After a relatively small initial SFT stage, most of Inkling’s post-training comes from asynchronous RL with more than 30 million rollouts. The desired effort is included in the system message, and the token length penalty is adjusted according to that value during RL. As previously discussed, a higher token cost encourages a shorter response. A lower token cost gives the model more room to reason. The table below summarizes what is actually documented in the six technical reports. Figure 32: Comparison of the disclosed training mechanisms and inference controls for six open-weight models with reasoning-effort settings. So, looking at the six different open-weight models, they have a shared framework. First, they introduce effort mode control through SFT and the chat template. Qwen3 explicitly mixes thinking and non-thinking examples, while GLM-5 adds interleaved, preserved, and turn-level thinking patterns. The second shared component is a mode-conditioned RL stage, where context windows and length penalties change with the requested effort. DeepSeek V4, Nemotron 3 Ultra, and Inkling use this approach. A third ingredient improves robustness under explicit budgets. Nemotron trains on randomly truncated traces, Qwen3 can continue from a forcibly stopped reasoning span, and Kimi alternates budgeted with unconstrained RL. These methods help preserve answer quality when the available reasoning length changes and is even cut short. The open-weight examples in this article implement reasoning effort through several different mechanisms. Similar labels can be backed by separate specialists, mixed SFT data, mode-conditioned rewards, hard token budgets, or combinations of these methods. It is difficult to say which approach is best. The models differ in their base checkpoints, training data, post-training compute, benchmarks, and serving goals. Their reports also omit many details needed for a controlled comparison. (Also, there may not be a one-size-fits-all, and a method that works well for an interactive assistant may be a poor fit for a long-running coding agent.) The holy grail is of course automatic effort selection. We saw this a while back with GPT 5’s Auto mode. It’s a tricky problem to solve, and in the end, the implementation was probably more miss than hit, which is why it got removed from the UI (at least, I can’t find it anymore). In the near future, I think reasoning effort will remain an explicit model input, which will most often be delivered through the system prompt. However agent wrapper/harness around the LLM, or an internal router may increasingly infer the appropriate mode and budget from the task state and available resources automatically (while of course still allowing a user override). I still hope that effort selection will become more automatic. Similar to GPT 5’s auto mode, a cheap model or router could choose the mode from the request, tool state, and remaining time or token budget while still allowing a user override. The override is useful if you want to optimize for latency or cost, or maximum performance. I realize that this was a long article, and it was perhaps not the flashiest topic. But I thought that given all the talk about LLMs, reasoning models, and agents, a look at reasoning models was something not covered before, and I hope it was a unique and somewhat useful overview! If you want a hands-on implementation of the core training methods behind reasoning models, my Build a Reasoning Model (From Scratch) book walks through reinforcement learning with verifiable rewards and inference-time scaling step by step, with code. This article focused on how a trained reasoning model can support different effort modes. The book takes a step back and shows how to turn a conventional LLM into a reasoning model in the first place. It is a sequel to Build a Large Language Model (From Scratch) and starts where that book leaves off. The print edition has now started shipping Build a Reasoning Model (From Scratch) [ Manning ] [ Amazon ] If you liked my previous Build a Large Language Model (From Scratch) book, this is essentially a sequel implementing inference-time scaling techniques and reinforcement learning algorithms from scratch. And if you want to support future long-form articles like this one, consider becoming a paid subscriber . It helps me keep writing these independent deep dives and sharing the accompanying code, figures, and experiments. Figure 1: The GPT 5.6 Sol model with different reasoning effort settings. (Benchmark numbers for Ultra are currently not available but should be relatively similar to Max, since it uses a similar effort level but accelerates the work with four subagents.) So yes, reasoning models are here to stay. They have become a standard part of modern model releases. In the past, I covered the methodology of reasoning models ( Understanding Reasoning LLMs ) as well as relevant research papers ( The State of Reinforcement Learning for LLM Reasoning and The State of LLM Reasoning Model Inference ). And I even wrote a whole new 440-page book on how to develop reasoning models, Build A Reasoning Model (From Scratch) . Figure 2: My new Build A Reasoning Model (From Scratch) book. In color! These resources have focused on turning a conventional LLM into a reasoning model. Now, in this article, I want to focus on and explain how to develop a reasoning model that has multiple effort modes, similar to what’s shown in the figure at the beginning of this article. No worries, this article can be read as a standalone article. However, the aforementioned resources may be interesting and useful. 1. A brief definition of reasoning models When talking about pretty much any machine learning or AI technique or subfield, the one lesson is that we usually shouldn’t take technical terms “literally”. For example, an (artificial) neural network in machine learning and AI doesn’t literally work like a biological neural network like the human brain. Similarly, when talking about “reasoning models”, we shouldn’t expect that these models literally reason like us humans. In the context of AI and LLM research, “reasoning model” means a model that outputs an intermediate reasoning trace, which is like an intermediate response that works through a question or task step by step. It’s probably easiest to explain this by showing an example. Figure 3: Illustration of a conventional LLM answer (left) and an answer by a reasoning model (right). 2. A brief overview of training and inference scaling reasoning models There are essentially two ways to improve (reasoning) task performance: training scaling and inference scaling. Figure 4: Training and inference-scaling are two ways to improve LLM and reasoning model problem-solving capabilities. Plot based on Learning to reason with LLMs Let’s briefly talk about training first. 2.1 Training reasoning models In a nutshell, DeepSeek-R1 proposed training an LLM using reinforcement learning with verifiable rewards (RLVR) to turn it into a reasoning model. RLVR is a technique to provide a reward signal ( and ) for verifiable data domains. These verifiable data domains here are math (we can use a symbolic math checker like SymPy or WolframAlpha to check results) and code (we can use a compiler or unit tests, or integrated platforms like LeetCode) to check for correctness. Figure 5: Illustration of accuracy and format rewards during RLVR training. Notably, the reasoning trace itself was not used for training or updating the model. Although they tried to use this intermediate response information for training, the DeepSeek-R1 paper reported that it wasn’t helpful for the model training, so it was ultimately not used. (Whether and how to incorporate intermediate reasoning traces in the training signal via process reward models is an active area of research.) Figure 6: The intermediate reasoning trace is ignored during RLVR; only the final answer and response format determine the reward. 2.2 “Aha” moments Anyway, just training on the output rewards alone, as Figure 7 shows, turned out to be sufficient for the model to learn how to reason through a problem, meaning that it would learn to write intermediate explanations, backtrack, and self-correct itself. These moments when the model realizes that it made a mistake and self-corrects itself are called “Aha” moments. Figure 7: An example of an aha moment, where a reasoning model notices an error in its intermediate reasoning and corrects it before producing the final answer. By the way, while DeepSeek-R1 is inarguably the more popular paper, and the paper that created excitement around reinforcement learning with verifiable rewards and the development of reasoning models, there is another paper, Kimi K1.5 , published on exactly the same day on arXiv (22 Jan 2025). Also, the term RLVR was already coined two months earlier in Tülu 3: Pushing Frontiers in Open Language Model Post-Training . ​ One reason why the DeepSeek R1 is ultimately the more popular paper is that it demonstrated that reasoning behavior can be achieved with pure reinforcement learning (RL). Figure 8: DeepSeek-R1-Zero applies RLVR directly to the pretrained base model without supervised fine-tuning. For instance, Tülu 3 and Kimi K1.5 applied reinforcement learning on top of a supervised fine-tuned (SFT) model. The DeepSeek-R1 model was also trained from an SFT checkpoint of the DeepSeek-V3 base model, and it included a DeepSeek-R1-Zero variant trained with pure RLVR. R1 Zero is a weaker model than R1, but it showed that RLVR is sufficient for teaching the model to generate and use reasoning traces. ​While R1-Zero was more of a proof-of-concept model, note that the full DeepSeek-R1 reasoning model training pipeline is usually multi-stage and a bit more complicated, as mentioned above. Figure 9: More detailed reasoning model training pipeline. This one depicts the various DeepSeek-R1 models. For more details, see my other article: Understanding Reasoning LLMs ​ By the way, most of today’s LLMs are effectively reasoning models, meaning they have been trained in a similar fashion to DeepSeek-R1 using a form of RLVR. 2.3 Inference scaling in a nutshell Next to improving reasoning behavior through training, another lever for improving model performance is inference compute scaling. In short, this means that we are spending more compute after training the model, during usage, to get better answers. This is a whole topic by itself, and you could read through my The State of LLM Reasoning Model Inference for a more detailed rundown: I will try to summarize what’s most essential to mention as background info below. First, training a model with RLVR is already implicitly leading to a form of inference scaling, since reasoning models usually output more tokens during inference compared to conventional LLMs, and that means we are spending more compute during inference. Second, we can further adjust this output length via reasoning effort levels, but more on that later. ​Third, there are many additional inference scaling techniques. A popular one is self-consistency, which is often implemented as a form of majority voting where the model is queried multiple times, and the final answer is selected via majority vote. Figure 10: An example of self-consistency, a popular inference scaling technique. This can be applied to conventional LLMs as well as reasoning models. Also, this method can be used on demand and in addition to reasoning training. A good example of that is DeepSeekMath-V2, where the researchers applied extreme inference-scaling on top of a reasoning model (specialized for math) to achieve state-of-the-art performance on challenging math olympiad-type problems. Figure 11: Two types of inference scaling (self-consistency and self-refinement) used together to improve math performance. Figure adapted from DeepSeekMath-V2: Towards Self-Verifiable Mathematical Reasoning But again, I will refer to my other article, The State of LLM Reasoning Model Inference for an overview of other techniques: 3. Think tokens You may have seen the tokens in the earlier “Aha moments” figure. I also included the corresponding figure below so you don’t have to scroll all the way up. ​ Figure 12: Common formatting tokens in reasoning models. These and tags are cosmetic with respect to reasoning ability. They do not make the model reason, and they are not required to achieve good reasoning performance. One could train the same model without these delimiters and likely reach similar benchmark performance. The purpose of these tags or tokens is mainly to mark where the reasoning trace begins and ends so that the training pipeline or user interface can separate it from the final answer and optionally hide it from the user. (UIs like ChatGPT or Codex usually do this.) The point here is that the tokens are not giving the model the ability to “think” or reason or reason better. One could train the same models without such tokens and reach similar benchmark performance. There is also nothing special about the literal strings and . Another pair of delimiters could serve the same purpose. By the way, the way this is implemented is typically by adding a formatting reward during the RLVR stage. So instead of just rewarding the model based on answer correctness, one would provide additional reward for the use of <think> tokens, which in turn encourages the model to use those. In DeepSeek-R1, for example, the overall reward was calculated as where the format reward was a simple rule-based check that encouraged the model to place its reasoning inside: reasoning trace . 4. Reasoning mode on and off switches The first generation of reasoning models was dedicated reasoning models. With that, I mean that there was a DeepSeek-V3 base model and a separate DeepSeek-R1 reasoning model. No matter what the prompt is, R1 generally outputs very verbose responses using lots of tokens, even for simple prompts. It also lacks a built-in option to turn off the reasoning mode. Figure 13: Reasoning models are very verbose, even for the simplest prompts. Later models, like Qwen3 and others, experimented with hybrid approaches, where the same model can behave like a regular instruction fine-tuned model or a reasoning model on demand. Note: Some model developers call this “thinking mode,” while others call it “reasoning mode.” Both terms refer to the same behavior. In Qwen3, this is handled via the tokenizer using or . Under the hood, setting essentially adds an empty section to the beginning of the assistant response to turn off Qwen3’s reasoning (”thinking”) mode. Figure 14: Response of Qwen3 0.6B reasoning model with and . (The empty tags are hidden in the interface on the left as they are part of the modified input prompt, not the generated answer.) How is this implemented during training, such that the model supports this toggle during inference time, as shown in the figure above? In short, as explained in the Qwen3 technical report , this on/off behavior is introduced primarily through supervised fine-tuning (SFT) and then reinforced during general RL in their largest flagship models. For instance, after the initial reasoning model is trained via long-chain-of-thought SFT and reasoning RL, they add a “Thinking Mode Fusion” stage. During this additional SFT stage, the model sees both thinking and non-thinking examples: Figure 15: “Thinking Mode Fusion” in Qwen3’s training pipeline to enable the reasoning mode on and off switch. In other words, the tokenizer does not add to the query. It directly fills in the empty section at the beginning of the assistant response. The model only sees the resulting tokens and continues directly with the answer. Anyway, this on-and-off toggle is essentially a simplified version of the reasoning effort levels in GPT-5.6 and others, which I cover in the next section. 5. How “reasoning effort” settings work In this section, I want to provide a brief overview of how the different reasoning effort toggles may be implemented, which have been introduced in models like GPT 5 and are present in pretty much any flagship model today. Concretely, at the beginning of this article, I showed a figure from the Codex GPT 5.6 interface that lets users select multiple reasoning “effort” settings. Figure 16: GPT-5.6 exposes six reasoning effort settings, ranging from Light to Ultra. The following subsection will illustrate how these settings may be implemented. Then, in the next section, I will go over some of the more interesting research papers related to this topic. 5.1 Reasoning effort and response length and quality Unfortunately, the implementation details of their effort settings are not shared by OpenAI, but there is some evidence out there that can be used for educated guesses. For instance, via their open-source gpt-oss models from last year (I wrote about them in From GPT-2 to gpt-oss: Analyzing the Architectural Advances ), we know that OpenAI allows us to toggle the reasoning effort setting via the system prompt (”Reasoning effort: low/medium/high”) that is prepended to each prompt. Figure 17: The gpt-oss chat template inserts the selected reasoning effort into the system message before sending the prompt to the same model. As expected, the reasoning effort directly affects the response length and accuracy, as shown below. Figure 18: Response length and quality of gpt-oss models under different reasoning efforts (annotated figure from the model card ) Presumably, their GPT 5 models, including the recent GPT 5.6 models, use a similar approach. By the way, note how different effort settings scale the response length in the figure above. The effort level seems directly correlated to token usage, which in turn seems correlated to accuracy. It might be possible to come up with effort settings beyond the “high” one, but I assume performance would saturate at some point. This saturation can be seen more clearly for the GPT 5.6 Sol model, which also shows that increasing reasoning budgets can become uneconomical at some point. Figure 19: Reasoning effort increases both API cost and coding-agent performance, with diminishing returns at the highest GPT-5.6 settings. Figure based on the Artificial Analysis Coding Agent Index v1.1. Another good, very recent data point that shows the relationship between reasoning effort, token usage, and benchmark performance is this week’s new open-weight Inkling release by Thinking Machine Labs. Figure 20: Increasing the Inkling effort level generally increases generated tokens and benchmark performance, with diminishing or uneven gains at higher effort. Figure from the Inkling announcement blog . As discussed in this section, during inference, the reasoning effort level can simply be controlled via a system prompt. (The ChatGPT UI presumably simply maps the menu choice to a system prompt.) However, this would not work for an arbitrary model and requires certain modifications to the training pipeline, which will be discussed next. 5.2 Possible effort level implementations While the training details are not public, neither for GPT 5.6 nor the open-source gpt-oss models, typically, the reasoning effort label is included in prompts during post-training. There are typically two ways to implement this. First, we can implement it as part of the RLVR process and apply a different length penalty when different system prompts are used. For example, a high length penalty when “Reasoning effort: low” and a mild or no penalty when “Reasoning effort: high”. Second, we can fine-tune the model after RLVR to follow different effort instructions via supervised fine-tuning (SFT). For instance, after the core RLVR stage and during SFT, the prompts in the training dataset are paired with target responses that exhibit the desired amount of reasoning. (The targets may be written by humans, generated by another model, or generated and then filtered.) Figure 21: Illustration of effort-conditioned RLVR and SFT. (This is a possible implementation, not a confirmed description of OpenAI’s training pipeline.) During this SFT stage, the model learns the association between the effort label and the target reasoning length directly from the training examples. An RL-based implementation would instead place the effort labels and budget-aware reward inside the RLVR stage. The two approaches could also be combined, which I suspect was done for both gpt-oss and GPT 5.6 (note that effort settings in GPT 5.6 are likely just changing the system prompt for a given user query). 5.3 Inkling case study The just-released Inkling technical report gives a small but somewhat concrete example of effort-level training. Figure 22: Inkling sweeps a continuous effort value between 0.2 and 0.99; higher effort generally produces longer responses and higher benchmark scores. During large-scale RL, they did two things for each sample: Specified the desired effort level in the system message. Adjusted the cost assigned to each generated token. Low effort uses a larger per-token cost, encouraging shorter reasoning traces. High effort uses a smaller per-token cost, allowing the model to spend more tokens. Figure 23: The model selection and reasoning effort menus correspond to two different scaling axes. Selecting Luna, Terra, or Sol changes the model, whereas changing the reasoning effort adjusts the inference-time compute for a fixed model. One small terminology caveat is that selecting a different model from the menu is not training scaling at that moment. The training has already happened. It is better to think of the model menu as selecting among models that were produced at different training scales. The Artificial Analysis results below show how these two axes interact in practice. Each blue curve corresponds to one model, Luna, Terra, or Sol. Moving along a curve by increasing the reasoning effort is inference scaling. Moving from one model curve to another corresponds to model scaling, which I use here as a practical proxy for training scaling. As expected, both approaches can improve the benchmark score, but they also increase the cost. More interestingly, the curves overlap. For instance, a smaller model at a higher reasoning effort can sometimes reach a similar score as a larger model at a lower reasoning effort. Figure 24: Training scaling and inference scaling for the GPT-5.6 model family on the Artificial Analysis Coding Agent Index. Moving along each model curve corresponds to increasing the reasoning effort. Moving across the Luna, Terra, and Sol curves corresponds to selecting a different model. By the way, the x-axis in this figure shows API cost rather than raw compute. The API cost is a useful practical measure, but it also depends on the provider’s pricing and the number of generated tokens. Also, the exact shape of these curves is benchmark-specific. So, the model size and reasoning effort form two separate knobs. We can use a larger model, increase the reasoning effort, or combine both. Which combination is best depends on the desired accuracy, cost, and latency. So far, the article should give you a pretty solid understanding of how reasoning effort modes work and how they are implemented. This is a fine point to wrap the article if you are short on time. Otherwise, if you want to look into some of the nitty-gritty details of some of the recent open-weight models, please read on! 6. Bonus: Different ways to implement reasoning efforts (in flagship open-weight LLMs) [This section is fine to skip unless you are interested in some additional details] Section 5 described two possible ways to train reasoning-effort controls, namely effort-conditioned supervised fine-tuning and reinforcement learning with different token costs. Originally, I wanted to cover research articles on alternative ways to implement reasoning budgets. However, reading through most of these articles, they seemed more like proofs-of-concept that may or may not work well in practice. So, instead of covering those, I decided to pivot a bit and cover those recipes used by state-of-the-art and notable open-weight (flagship) LLMs. For these models, there is at least evidence that the methods work in practice. This leaves six examples. DeepSeek V4, Nemotron 3 Ultra, Kimi K2.5, GLM-5, Qwen3, and Inkling. They have different levels of detail in their reporting, but each contributes a useful variation. (I exclude models whose reports only show an effort setting in the user interface without explaining how that behavior was trained.) 6.1 DeepSeek V4 trains separate effort specialists Let’s start with the DeepSeek V4 technical report , which describes the use of three modes: Non-think produces a direct response without a reasoning trace. Think High is the classic approach where the model places the reasoning trace between <think> and </think> tags. This is similar to what was discussed in the DeepSeek R1 section (section 2) at the beginning of this article. Think Max is the same as above but adds a special system instruction. (More on that below.) Figure 25: Reasoning effort control overview from the DeepSeek V4 documentation At first, this sounds like a simple prompt engineering trick, but this prompt is actually backed by a different training setup. That is, each mode uses its own context window and length penalty (unfortunately, the exact length penalty implementation is not detailed in this report). Think Max receives a longer context window and a smaller length penalty than Think High, which gives it more room to continue reasoning. So, the system instruction selects a behavior that was created during post-training. Adding the same instruction to an arbitrary model would not have the same effect. Figure 26: DeepSeek V4 describes the three effort modes and the larger teacher pool in separate parts of the report. The teacher pool contains more than ten domain specialists. The report does not disclose how these teachers map to Non-think, Think High, and Think Max. Unfortunately, the public, and otherwise very detailed DeepSeek V4 report does not connect the descriptions of the reasoning mode and domain specialists in enough detail to reconstruct the exact teacher assignment. However, the report states that the final model, which supports different reasoning effort levels, was created via on-policy distillation from said teachers. To summarize, DeepSeek V4 develops the three reasoning specialists during post-training. Starting from the base model, it applies supervised fine-tuning followed by RLVR via GRPO. The RL configuration differs for each mode. In particular, each specialist uses its own context window and length penalty, while Think Max additionally receives a special system instruction. Then, including domain specialists, the different reasoning mode specialists are distilled into a single checkpoint that supports all three effort modes. 6.2 Nemotron 3 Ultra combines learned modes with hard budgets The Nemotron 3 Ultra technical report describes three settings called reasoning-off, regular, and medium-effort, analogous to DeepSeek V4 in the previous section. Medium-effort is the cheaper reasoning mode compared with regular. NVIDIA introduces this mode during SFT using examples generated by GPT-OSS-120B in its medium-effort mode, and then further optimizes it during RLVR. About 2.5% of the RLVR prompts use medium-effort (this corresponds to length-based adjustments applied to their rewards). 6.2.1 Using Nemotron reasoning budgets during inference At inference time, all three modes are selected through the chat template . Figure 27: Nemotron 3 Ultra reasoning settings via the chat template (examples from the official model card ) 1) Regular is the default and uses , which starts the assistant response with an opening tag. 2) Medium-effort uses together with where the latter setting also appends {reasoning effort: efficient} to the latest user message. By the way, to further complicate things, the regular and medium-effort modes can also be combined with a separate inference-time reasoning budget. This budget acts as an external stopping mechanism. In the released implementation, the chat client asks the model to end the reasoning trace near the chosen token limit. If the model has not emitted , the client closes the reasoning block and continues generation to produce the final answer. The learned effort mode determines how the model uses its reasoning tokens, while the budget constrains how long the reasoning trace can continue. This makes it possible to pair either mode with a tighter or looser budget depending on the desired cost and accuracy. 3) Reasoning-off uses , which prefills an empty block (similar to Qwen3 discussed in section 4) so that the model proceeds directly to the final response. Thus, these are chat-template controls rather than system prompts. 6.2.2 Reasoning budget-aware training in Nemotron The inference controls described above are backed by two related SFT components. The first introduces medium-effort behavior using GPT-OSS-120B traces, as discussed earlier. The second prepares the model for hard reasoning budgets. To construct this training data, the authors take regular reasoning traces, truncate them at randomly selected token budgets, and keep the original final answers. The inserted token is masked from the SFT loss. As a result, the model sees examples where it has to move from an incomplete reasoning trace to the answer after the reasoning block has been closed externally. Medium-effort training then continues during RLVR. About 2.5% of the RL prompts use the medium-effort setting across math, STEM, and coding tasks. The report notes that the mode can be calibrated through reward hyperparameters where length-based reward adjustments provide additional control over the cost-quality trade-off. Figure 28: Nemotron 3 Ultra introduces medium effort with teacher-generated SFT data, random-budget truncation, and a small medium-effort subset during RLVR. 6.3 Kimi K2.5 alternates budgeted and unconstrained RL The Kimi K2.5 technical report discusses a training method called Token Efficient RL for lower reasoning effort. (While there was a K3 announcement this week, the reasoning-effort methodology of K3 is not publicly disclosed, but it could be similar or related to K2.5.) 6.3.1 Kimi’s Toggle method The report mentions that a fixed token budget can make a reasoning model overfit to short solutions. That means the model becomes more concise (i.e., faster and cheaper), but it may lose the ability to benefit from additional inference-time compute and can thus perform poorly. Figure 29: The proposed Toggle method makes Kimi K2.5 much more token-efficient while keeping the overall benchmark performance similar. Annotated figure from https://arxiv.org/abs/2602.02276 Kimi K2.5’s method, called Toggle, alternates between two RL phases every fixed number of training iterations: 1. In the budgeted phase, correct solutions are encouraged to stay within a problem-specific token budget. 2. In the unconstrained phase, the usual maximum generation length is restored so that the model can still learn from longer solutions. For each problem, the budget is estimated from a selected percentile of response lengths among correct rollouts in RLVR. The budget constraint is then only activated once the mean accuracy on that problem exceeds a threshold. This avoids forcing the model to shorten its reasoning before it can solve the problem reliably. Figure 30: Overview of the two phases of the Toggle method. The report evaluates Toggle on K2 Thinking and finds that it reduces generated tokens by about 25 to 30% with little change in benchmark performance. The same behavior also transfers from math and coding RL tasks to GPQA and MMLU-Pro. Toggle supplies a concrete flagship-model recipe for training a more token-efficient reasoning policy while preserving its ability to scale at test time. 6.3.2 What Toggle changes at inference Toggle operates entirely during RL training. Both alternating phases update the same policy (i.e., LLM), and the final (unified) checkpoint has no budgeted-versus-unconstrained selector. At inference, the resulting model then runs in thinking mode by default. Interestingly, though, Kimi K2.5 itself exposes a separate binary choice between thinking and instant modes in some APIs I checked (like vLLM or SGLang). Thinking mode is enabled by default. Instant mode disables the reasoning trace through thinking: in the official API or when serving the model through vLLM or SGLang. However, these settings are separate from Toggle. Also, the official Kimi report does not provide a separate training recipe for instant mode. However, K2.5’s SFT data were generated using both the earlier K2 model, which produces direct responses without long reasoning, and K2 Thinking, which produces extended reasoning traces. This likely exposes the unified checkpoint to both response formats similar to what’s done in Nemotron 3 above. At inference time, the chat template selects between them by prefilling either an open <think> tag for thinking mode or an empty block for instant mode. But again, unfortunately, the report does not disclose the exact data mixture or whether additional mode-specific RL was used. The newer Kimi K3 provides a more direct inference-time effort interface. The current Kimi Code documentation lists three settings called low, high, and max, with max as the default. These are passed through the parameter. However, Moonshot has not yet explained how the three effort levels were created during training. Its launch post says that these details will appear in a future K3 technical report, so I’ll stay tuned for that. 6.4 GLM-5 introduces turn-level and interleaved thinking through SFT The GLM-5 technical report extends the binary on/off thinking switch introduced with GLM-4.5 to multi-turn and tool-using scenarios. It describes three related behaviors (rather than three effort levels): Interleaved thinking: this inserts a reasoning block before each response and tool call. Preserved thinking: here, the chat retains earlier reasoning blocks across turns so that the model can reuse them later. Turn-level thinking: this enables or disables reasoning separately for each request in a conversation.

0 views