Latest Posts (20 found)
iDiallo Today

You Can Drop SEO

I started learning web development around the time the term search engine optimization, SEO, was becoming more common. You could still see the shift between pre-SEO content and post-SEO content. A popular article titled "Forgotten" (a great story) was suddenly republished as "Forgotten: My Adventures as an Employee the System Forgot to Erase." Both titles and content were filled with keywords that signaled to search engines what the page was about. Articles basically catered to the search engines and only incidentally served people. On my own website, I remember editing page titles several times, then waiting a few days to see if Google noticed the changes and ranked me better. In fact, before Google introduced personalized results, I built a tool at my job to track how our website ranked for different combinations of keywords. When personalized search results arrived, a lot of tools became obsolete. It was still useful, though, to track keywords for generalized metrics. But now AI Overviews are here. And it's not just Google's, plenty of people go straight to ChatGPT and ask the large language model their questions directly. While the information returned might be sourced from real websites, there's almost no incentive to click through to those sources. While my inbox is still flooded with people promising to improve my SEO, I think it might not be helpful at all anymore. My traffic has shifted from mostly coming from Google to just a handful of visits. Most traffic now comes from AI bots scraping my content, and RSS readers (thank you!). I take this as a relief. We can finally drop the act. We don't have to write keyword-stuffed titles and blog posts just to appear in search results. Large language models can understand our content just fine without it, and we won't be getting that traffic anyway. We can drop SEO. You can finally write for yourself. Write for your audience. No need to cater to the robots.

0 views

2026.37: Duo Threats

Welcome back to This Week in Stratechery! As a reminder, each week, every Friday, we’re sending out this overview of content in the Stratechery bundle; highlighted links are free for everyone . Additionally, you have complete control over what we send to you. If you don’t want to receive This Week in Stratechery emails (there is no podcast), please uncheck the box in your delivery settings . On that note, here were a few of our favorites this week. This week’s Stratechery video is on Autonomy and Innovation . The Duo Arrives.  Does the world need a foldable iPhone that costs between $2000 and $3,200? It’s a fair question. On the other hand, life is short, and tech is a lot more fun when we have new and possibly-crazy hardware projects to discuss — particularly when they’re deployed by Apple. To that end, I heartily recommend cleansing your palate from a week of media-wide AI angst by reading Ben’s take on Apple’s iPhone event and pairing that with Friday’s Dithering , and John Gruber’s impressions of Duo-mania on the ground in Cupertino. Also: bonus being-right points to Gruber, who nailed the name of this device back in April .  — Andrew Sharp AI That Benefits Humanity. I loved Wednesday’s Update contrasting OpenAI’s thrilling and technically impressive Navier-Stokes breakthrough with the release of Meta’s far less sexy Muse agent. While OpenAI’s tactics may in fact chill research in advanced mathematics, what Meta has assembled is free (to consumers) hardware and software that dramatically reduces the barrier to entry for ordinary people looking to harness the power of agents, making the AI upside a lot more accessible to the masses who don’t want to buy a Mac Mini. That’s a big deal! We discussed Muse more on this week’s Sharp Tech , including tips for getting started with agents, and questions about whether people will actually take advantage of this opportunity.   — A S Closing the Book on a Catastrophe.  Everyone’s familiar with the benefits of pro sports ownership and its ability to turn semi-anonymous rich guys into full blown celebrities, but Microsoft co-founder Steve Ballmer is now a living testament to the unstated risk — sports ownership fame can, in a worst case scenario, become infamy. Last week his Clippers received the harshest penalty in NBA history for circumventing the salary cap to pay Kawhi Leonard. We recapped all of it on GOAT this week , including successes and failures in sports journalism, why Kawhi got off easy, and the staggering amounts of evidence that sealed Ballmer and the Clippers’ fate.  — AS Write Things Down — Writing things down is powerful, for humans and for AI; what comes first, however, is what to write, why to do it, and actually getting things done. OpenAI Does Math, Reward-Hacking, Meta Launches Personal Agent — OpenAI solving one of the most famous math problems is extremely impressive, and of little impact to most people’s lives; Meta’s Muse agent launch has the potential to be the exact opposite. The iPhone Duo, The Intelligent Personal Hub, Apple Watch Audio Intelligence — Apple once again demonstrated the power of integrating hardware and software, but it’s biggest AI blindspot might be its belief in the primacy of apps. Agents and Forklifts The Flood that Wrecked the Hard Disk Drive Industry Did Numerical Control De-skill Machinists? Closing the Book on the Clippers Catastrophe and Early Over/Under Picks in the Atlantic Astra (and AGI?) Arrives, Meta’s Muse and the Agent Opportunity, Anthropic and the Revival of (P)Doom Angst

0 views
Unsung Today

Key symbols we lost to time, pt. 1: The PC side

Various old computers had their keyboards adorned with unique symbols. Companies like Commodore , Atari , Amiga , or even – in its previous life – Apple chose to put their company logos on keys, and there were other weird and obscure keys on weird and obscure keyboards. But it was Apple’s recent push to move their American keyboards closer to European ones by embracing more iconography, that made me think of forgotten key symbols less obscure, ones that belonged to platforms we still use today. Even on a Mac and a PC, some key symbols didn’t make it to modern times. So let’s start with the PC side today since that part of the story begins earlier, and do Macs in a follow-up post. For a lot of 20th century, a battle has been waging between words and icons. The first salvo was, perhaps, the traffic signs : America embraced words, while Europe relied more on iconography. (As much as it looks like it, it wasn’t just “graphic design vs. not”; as a more varied continent with multiple languages, Europe needed a more universal visual language to help people travelling between countries.) This, I understand, trickled down to other things: home electronics, and computers. There, iconography also made it easier to make one product and sell it across all of Europe, without needing to introduce many SKUs with different UI strings. Here’s IBM’s Selectric typewriter from the 1970s, in its American and European edition: (If you’re curious, Express was a very fast Backspace, and Index moved the page down; both were prototypes of future arrow keys.) Here’s IBM’s early 1130 computer from 1965, which sported an unusual symbol for space: Some IBM laboratory and scientific computers in the 1970s and even 1980s veered more into iconography, but eventually lost to text as office PC users rejected the confusing symbols. As their keyboards morphed into PC/Windows keyboards we know today, only four symbols remained and gained widespread acceptance: ⇧ for Shift, ↵ for Enter, ⇥ for Tab, and some version of an arrow for Backspace. But let’s look at those old symbols, some beautiful, all interesting. The two symbols below are: Print Screen (old CRT screen turning into a piece of paper) and key beep – popular when people were transitioning from loud typewriters to relatively quiet keyboards: = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/key-symbols-we-lost-to-time-pt-1-the-pc-side/5.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/key-symbols-we-lost-to-time-pt-1-the-pc-side/5.1600w.avif" type="image/avif"> = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/key-symbols-we-lost-to-time-pt-1-the-pc-side/6.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/key-symbols-we-lost-to-time-pt-1-the-pc-side/6.1600w.avif" type="image/avif"> Here – on the front edge of the also-forgotten Reverse Tab – you can see Home, which historically meant “return to the top left corner of the screen” and sometimes even “clear the screen”: = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/key-symbols-we-lost-to-time-pt-1-the-pc-side/7.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/key-symbols-we-lost-to-time-pt-1-the-pc-side/7.1600w.avif" type="image/avif"> But my favourites were these, for Insert (now gone) and Delete (still with us): = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/key-symbols-we-lost-to-time-pt-1-the-pc-side/8.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/key-symbols-we-lost-to-time-pt-1-the-pc-side/8.1600w.avif" type="image/avif"> = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/key-symbols-we-lost-to-time-pt-1-the-pc-side/9.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/key-symbols-we-lost-to-time-pt-1-the-pc-side/9.1600w.avif" type="image/avif"> These seem inspired by proofreader marks, which feels wonderfully old-time’y: = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/key-symbols-we-lost-to-time-pt-1-the-pc-side/10.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/key-symbols-we-lost-to-time-pt-1-the-pc-side/10.1600w.avif" type="image/avif"> = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/key-symbols-we-lost-to-time-pt-1-the-pc-side/11.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/key-symbols-we-lost-to-time-pt-1-the-pc-side/11.1600w.avif" type="image/avif"> Building on that visual language, one could also find invert/​reverse video, blinking, and underline: = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/key-symbols-we-lost-to-time-pt-1-the-pc-side/12.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/key-symbols-we-lost-to-time-pt-1-the-pc-side/12.1600w.avif" type="image/avif"> And this absolute beauty, which I think meant “delete word”: = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/key-symbols-we-lost-to-time-pt-1-the-pc-side/13.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/key-symbols-we-lost-to-time-pt-1-the-pc-side/13.1600w.avif" type="image/avif"> The really interesting thing is that some of those symbols survive today in Unicode. I spotted at least ⎀, ⎃, ⎁, and ⎂. The last two are for contiguous and non-contiguous underline, which I feel is a story I should know, but I don’t (yet).

0 views
Unsung Today

Fingers don’t look around

Buttondown, a newsletter publisher, has a pretty standard CSS editor in its web app. You can edit the code and whenever you make any change, you can then press the Save button to make it go live: The view also thoughtfully supports pressing ⌘S to do the same thing: However, if there is nothing to save, ⌘S is ignored by the CSS editor, and falls back to the browser’s handling: It feels logical: ⌘S only takes effect when the button is visible, otherwise why would a user press it? In a front-end sense, it might even seem thoughtful. We all witnessed web apps that greedily took over some interaction, and broke things in the process. Hell, I did that myself . In theory, you – the user – take a careful look at the state of things, notice the button, and then press ⌘S. But in practice? It’s none of the above. Fingers don’t look around. You might press ⌘S twice in a row. Or after an undo to an already saved state. Or after pressing another key so light it didn’t register. Or after just sitting down to an open document that’s already saved. Or because you weren’t sure if the previous ⌘S press worked. Or just to be sure. You might not know why, and you might not even notice. The beautiful raw power of ⌘S as a citizen of motor memory is that it’s automatic, mindless, habitual . That’s why ⌘S here needs to be both deterministic and idempotent – if there is nothing to save, don’t let it fall back to the browser, don’t show an error message, don’t beep at the user. Just ignore the keystroke altogether. It might feel funny, but there are tons of places in the UI that already look the other way. Off the top of my head: It gets a lot more interesting than these, but we’ll talk about more examples of “finger logic” in future posts. if you center align already center-aligned text in any writing app, the app just ignores you, if you press ⌘A to select all more than once, no one’s shouting at you, if you try to click a disabled button, the click gets quietly swallowed.

0 views

Premium: The Hater's Guide To Broadcom

According to The Information , in early 2024, Broadcom CEO Hock Tan hosted a “coffee chat” with employees of the recently-acquired VMWare , and introduced them to his particular brand of management: He may not be your dad, but Hock Tan sure is a motherfucker. Broadcom is a company you likely know for its XPU platform — a collection of different bits of intellectual property and access to semiconductor parts that allow it to build custom AI chips, the best-known of which are Google’s TPUs . It just signed a $30 billion deal with Apple to build “custom ASIC silicon products .”  Apple was already a massive customer of Broadcom, which historically provided a good chunk of  the wireless and radio frequency parts that you’d find inside iPhones and its other devices, representing at one point more than 20% of revenues, dropping to around 10% to 15% with the growth of AI chip sales and the acquisition of VMWare. For the most part, Broadcom’s business is built on selling companies the internal bits and pieces of either their hardware or the hardware surrounding their hardware — everything from wireless and RF components to data center networking tools.  It also dabbles in mainframe software (from its acquisition of CA Technologies ), security (from its acquisition of Symantec’s enterprise security business ), and virtualization software (from its acquisition of VMWare), and these segments cost it a combined $99.1 billion in cash and stock (not counting for inflation).  Except “Broadcom,” as a company, wasn’t always called Broadcom, and wasn’t founded by Hock Tan. As I’ll get into in this piece, “Broadcom” was once two very different companies — a wireless communication chips company founded in 1998 called “Broadcom,” and the private equity-formed monstrosity formed out of a spun-off semiconductor subsidiary of Hewlett Packard called “Avago Technologies.”  Much like Oracle , Broadcom is the story of a company acquiring other companies and then screwing over both its customers and employees in the pursuit of endless growth, which usually involves price gouging, massive layoffs, and cost-cutting anywhere that won’t improve margins. A great example came from a Wall Street Journal piece from January 2018 involving Broadcom’s failed $117 billion attempt to acquire Qualcomm:  Two months later, the deal would collapse despite a dozen banks signing on ( per Reuters ) to provide Broadcom a $100 billion bridge loan to get the deal done, with President Trump vetoing the deal to avoid Broadcom (then a Singapore-based company) exercising control over the US-based Qualcomm. To give some credit to the administration, the CFIUS had ( per The Hill ) “...worried that Broadcom’s takeover would lead to a decline in investments in research and development in the sector, opening the door for Chinese firms to take the lead in developing next-generation wireless technology.”  That R&D point was a very real concern. Per The Journal: Pffft, 19%? That’s chump change. Since the acquisition of VMware, Broadcom’s R&D budget as a share of revenue has decayed to an unremarkable 9.8% of revenue in its latest quarter.  On a trailing-twelve-month basis, Broadcom is exceptional among its peers for how little it invests in R&D as a percentage of revenue, beaten only by NVIDIA, which has the excuse that it is the literal largest and most-profitable company on the US stock market.  That’s because Broadcom doesn’t really do “innovation” or care about “being good to its customers, but by hoarding other people’s patents, iterating on their creations as little as necessary, and making it impossible to avoid wiring Hock Tan money. Even its FBAR filters (used to block out interference on mobile phones) — a critical part of its deal with Apple — come from Avago’s acquisition of the original Broadcom. A year or two ago, this could’ve been called The Hater’s Guide To Avago , because that’s really been the story of Broadcom — a Singaporean semiconductor firm that rolls up other companies’ technology under a brand made famous by somebody else.  Per The Wall Street Journal, a few months before its acquisition of Broadcom :  Much like Oracle , Broadcom used M&A as a means of treading water revenue-wise, with each one having little effect on its overall trajectory outside of its acquisition of VMWare. Yet Broadcom had been building something quietly behind the scenes through the combined acquisitions of LSI (which had merged with Agere a few years previously) and its own semiconductor might — a budding relationship with Google to build its Tensor Processing Units (TPUs), AI chips that at first worked to support products like Search and Maps, and would eventually become a huge part of the AI boom.  To be clear, Broadcom didn’t “see anything coming” or “catch the AI boom in its infancy.” While it deserves some credit for rolling up various different semiconductor companies like it’s playing Katamari Damacy, this is not a situation where Hock Tan or anyone had any kind of precognitive event that made them invest in ASICs in anticipation of a massive payoff.  What actually happened was far simpler: Google, which had already been running its services powered by (non-LLM) AI, got Broadcom to use its pile of various patents and supply chain connections to put together specialized silicon that had incremental boosts to Broadcom’s revenues until the launch of ChatGPT scared Sundar Pichai into sinking billions — and then tens of billions — of dollars into successive generations of TPUs. And in Fiscal Year 2024, Broadcom began breaking out that revenue from its semiconductor solutions division, and something became alarmingly clear: that it’s become dependent on AI revenues for virtually all of its future growth. Between Q1 FY2024 and Q3 FY2026, AI revenue has gone from 19.2% to 56.4% of Broadcom’s revenue, with analysts expecting it to make up 68.4% of FY27, 79.8% of FY28 and 81.9% of FY29. And you’ll never guess who the customers are!  That’s right — OpenAI (for its Jalapeno AI chip) and Anthropic (buying Google TPUs) , who are set to become Broadcom’s largest customers in Fiscal Year 2027, which means that tens of billions (and eventually hundreds of billions) of dollars of revenue will be tied to whether two unprofitable, unsustainable AI companies can afford to pay.  This is the story of how a grab-bag of other people’s innovations has accelerated in the space of three years to become one of the largest AI chipmakers of the world, and how its desperation for growth has forced it to engage in the darkest forms of circular financing. This is The Hater’s Guide To Broadcom — the hard numbers and charts behind Hock Tan’s aggressive play to beat NVIDIA and become Google, Anthropic, and OpenAI’s chipmaker of choice…and how dangerous it might be if it fails.

0 views

2026-09-11 15:02: Because a basic sitemap would be boring! https://kevquirk.com/sitemap

Because a basic sitemap would be boring! https://kevquirk.com/sitemap Thanks for reading this post via RSS. RSS is ace, and so are you. ❤️ You can reply to this post by email , or leave a comment .

0 views

Ironman Training Diary - Learning How to Swim at 51

Channeling my inner Michael Phelps at 51 years old has been a humbling experience but I'm getting better.

0 views
Jim Nielsen Yesterday

An Ode to Links

A URL is a technical thing: a hostname, a path, a port, an origin, some query parameters, etc. But a link is a cultural thing: an invitation, a citation, a gift, a shortcut, a connection. All from a simple idea: here’s a thing that points to another thing, off you go! We’ve built an entire culture around links. Just look at the language we use for them: Or even just, “Link?” Our language spans the spectrum of meaning, from one end to the other: We make links. Links make the web. The web makes Us. Reply via: Email · Mastodon · Bluesky “Drop me a link.” “Send me the link.” “Got a link?” “Link in bio.” “Link in the description.” “Here’s the link.” “Don’t click that link.” “Sorry, wrong link.” We save links. We lose links. We share links. We hide links. We send links. We receive links. We collect links. We discard links. We open links. We close links. We add links. We remove links. We make links. We break links. We shorten links. We expand links. We follow links. We ignore links. We click links. We don’t click links. We link in. We link out.

0 views
Giles's blog Yesterday

Extending Raschka's GPT-2: an MoE trained from scratch on an RTX 3090

Mixture-of-experts models are really nifty. You get inference speed close to a small model's, with a lot of the smarts and knowledge of a large one. While they use as much memory as an equivalently-sized dense (non-MoE) model, they're much faster. The frontier labs don't publish their architectures, but Claude and ChatGPT are widely rumoured to be MoEs these days -- and certainly many large open-weights models like DeepSeek and Kimi K3 are. In this post, I'll show you how I added MoE support to the GPT-2-style code from Sebastian Raschka 's book " Build a Large Language Model (from Scratch) ", then used that to train a 446M-parameter model with 220M active parameters from scratch on my RTX 3090 -- essentially, GPT-2 small with 6 experts, 2 active per token. I wanted to understand how MoEs work, and as always, felt that the best way to do that is to build one (and then to write it up like this). Hopefully because it's all fresh in my mind as I write this, I should be able to explain things clearly for others who've finished Raschka's book. The training run took just less than eight days, and the resulting model got a better loss on my test set than any of the other models I've trained so far -- which was a good thing, given that it took four times longer to train! It was also better than the original OpenAI GPT-2 small (124M parameters), but not as good as GPT-2 medium (345M parameters), although it was close. That second result was interesting, as my model had more total parameters than GPT-2 medium, but fewer active ones. On an instruction fine-tuning test, it did better than any of my other models, but worse than both OpenAI models (why OpenAI's models are so good on that particular test is a mystery I'm digging into separately). I'll go into those numbers in more depth later on. Firstly, though, I'll run through the code that I needed to add to GPT-2 to make this work -- not just the code for the experts themselves, but the additional code for training it. With MoEs you can't just train to minimise loss on your training set -- you need to add on an "auxiliary" loss to make sure they actually use all of their experts, and that was where it became most interesting. Before we get into the weeds, though, let's start off with the basics; how do MoEs work in Transformers-based LLMs? The phrase "mixture of experts" evokes an idea of a model which has separate parts that are knowledgeable in different domains. Maybe one part would know about coding, another about history, another about philosophy, and so on. You can imagine something that had separate LLMs for different topics, and routed incoming inputs appropriately. That's actually not a bad design for a system -- Sakana.ai got a lot of interest for their Fugu system back in June of this year, and it works rather like that. But the "experts" in an MoE are at a much lower level, and (as with so many things in LLMs) their expertise is in some weird and alien thing that they determined was helpful for modelling language during their training. Let's look at how they fit in mechanically. The GPT-2-style LLM that we will start from -- the one Raschka describes in his book -- looks like this: Diagram 1 : a GPT-2-style LLM at the top level The MoE magic happens inside those Transformers layers, so let's zoom in on one of them: Diagram 2 : a GPT-2-style Transformers block Specifically, what we do to make our LLM an MoE is to replace that feed-forward network (FFN) with multiple separate FFNs -- our experts. Different context vectors get routed to different experts based on their contents. The FFNs themselves -- in both the dense and MoE versions -- are surprisingly simple . In GPT-2, they take in the incoming context vectors, run them through a normal linear layer to expand the number of dimensions by four, run the result through a GELU activation function, then project it back into the original incoming dimensionality with another linear layer. It's a really simple two-layer neural network, and at first glance seems somewhat arbitrary. Tutorials about LLMs tend to spend pages and pages explaining attention mechanisms, and pretty much gloss over the FFNs. But the FFNs take up twice as many parameters as the attention mechanism in GPT-2. They're clearly highly important, and my (very loose) metaphor for why that is, is that attention is how the LLM works out what to think about, and the FFNs are where it does its thinking. It's not a perfect match for what's going on, but I think it's a decent working model for intuition. An MoE leverages this. Instead of having just one FFN per Transformers block, we have multiple, and we use a subset of them for each context vector. We have what is called a router, or a gating network. It takes the incoming context vectors, and for each one, decides which of these FFNs -- which experts -- to use. Then we feed the input into the experts that were selected for it, combine their results, and that's our final output -- like this: Diagram 3 : an outline of an MoE block Doing this gains us more "space" in the LLM for it to remember facts and ways of thinking about things -- we have multiple experts for that knowledge to be spread over. We could, of course, do that by dedicating more space to the FFNs -- for example, by having one bigger one. But with an MoE, because we only activate a subset of the experts for each context vector, we save on the amount of computing we do for each context vector. We need to keep all of the experts in RAM -- remember that we're routing to them per-context vector, so in a batch of sequences we're likely to be using most if not all of them. But we don't have to feed everything through all of them. So, that's the basics -- nothing conceptually difficult. What becomes more tricky is the implementation. How, concretely, does the router choose which experts to use for a given context vector, how do we implement that choice -- and how can we do all of that efficiently? And how do we train the router to make its choices? When I started on this project, my first step was a Google search for useful references. I came across this excellent summary from IBM. In that post, they mention a number of papers; four seemed relevant 1 : I decided to take a look at them, and was pleasantly surprised about how readable they all were. It looked to me like I'd be able to put together a working MoE model by following what they (and the IBM summary) said, and that turned out to be true -- the model worked, and trained well. Of the four papers, I found the Switch Transformers one the most useful, but the original 1991 paper also clarified a lot to me. I really do recommend skimming through them if you want to learn more background about this stuff. Anyway, once I'd got it all put together, and trained the model, I compared what I'd written to the Hugging Face source code for Mixtral , which was the first open MoE model I remember hearing about. Mixtral has a bunch of other improvements over GPT-2 beyond its MoE support (RMSNorm, RoPE, and so on), and it's a much larger model than mine. But it turns out that the specific way it handles the MoE side of things is largely the same as mine, apart from one difference in how it calculates auxiliary loss (about which more later). So that was reassuring. I also ran the final codebase past three LLMs -- ChatGPT, Claude, and Kimi K3 -- just to make sure that I hadn't drifted from the mainstream or screwed up in other ways. I did this in an anonymous/private session so that they wouldn't use any memories they have of conversations we'd had in the past, and asked them Please take a look at the attached and tell me what you see. I'm particularly interested in the MoE stuff. All of them came back with some variation of "it's a pretty standard MoE implementation on top of GPT-2, with load-balancing adapted from Switch Transformers". Even more reassuring! So I'm comfortable that what I've built is a normal implementation, and that's what I'll describe in the rest of this post. It's all fresh in my mind, so hopefully by describing it in terms that would have made sense to me when I just finished Raschka's book, the result will be useful to other people coming to this for the first time. Let's get started by digging into the mechanics of how the router works. The problem we want to solve is that we have a context vector coming into our MoE block -- as per the diagram above -- and we want to decide which experts we want to route it to. Let's say that we have n total experts, and we want to send each input to k of them (where 0 < k < n , of course). We'll also say that our context vectors are of dimensionality d emb . Taking a d emb -sized input and categorising it into probabilities across a number of options is a pretty standard job for a neural network. In fact, we can use a single linear layer for it. Imagine that we create one with d emb inputs and n outputs. For a given context vector, it might produce (for n = 6 ) something like this: So we can imagine picking out the indexes of the top k of those. Let's be specific, and say that k = 2 . That means that we have indexes 4 and 2 -- the positions of the two largest numbers -- so we want to route our original context vector to experts 4 and 2. They do their calculations, we get output context vectors from them, and we can combine them. How might we combine them? Well, we do a lot of adding context vectors together in our GPT-2-style LLM code, and treat the results as meaningful -- token embeddings get added to position embeddings, shortcuts around attention and the FFN get added back in to their results, and so on. So perhaps we could do that? That kind of setup has a problem, though -- it doesn't train the router. Let's think about the MoE block again; here's the diagram again: Diagram 3 (repeated): an outline of an MoE block Think about the flow of data through it. The context vectors flow into the router. We use the output of the router to select the two experts, and then the original context vectors -- the ones that were fed into the router -- flow into those experts, then are combined, and we get our result. Now, consider what happens when you're training a model. You run some training data through it, then calculate the loss -- how good your model's results for that data are. You then use that to work out the gradients that you want to apply to the parameters to make it better. You work out the gradients by using back-propagation, working back from the loss, through the network, retracing the computation graph in reverse. When your backprop gets to this part of the model, it will start with the output context vectors, trace back through the combination step, then back through the two chosen experts, then back to the input context vectors -- and then it will go back to whatever step came before the MoE block. The calculations inside the router that selected our two experts did actually happen in our forward pass -- but they're not in the computation graph as we trace it backwards from the loss. It's kind of a dead-end. There's nothing in there to connect our selection of which experts we used to the path through the computation that ended up at the loss. (Interestingly, the effort of doing that diagram made that particularly clear to me -- that slightly-messy labelling of the arrows at the top, with numbers to show the sequence, is a direct result of the same issue.) What that all means is that the router will not be trained. The backward pass of our training will completely ignore it, and so there will be no gradients to apply to it. A starting model with random weights will randomly allocate context vectors to experts, and it will continue to do that. Clearly, we need to somehow modify our computation graph so that the router is connected -- so that when the backward pass flows up from the output context vectors, it sees the router -- and ideally sees some aspect of it that will allow it to be trained to do its job better. "Adaptive Mixtures of Local Experts" starts by describing some previous work which treats the router's outputs as weights for each of the experts. That's a nice trick, and while it does have some problems (as we'll see in a bit), it fixes the backprop issue. (Interestingly, they actually decided to do things differently -- but later work, like Switch Transformers, goes back to doing things this way, with some important tweaks.) As I said earlier, back in 1991 they weren't thinking of MoEs as being a way to save on computation. Instead, their focus was on training better models to handle specific tasks, and they felt that having some way to split those tasks up might make that easier. So, for their case, they might take those outputs from above and normalise them: ...and then run their inputs through all six experts, then add together the outputs weighted by those numbers -- 0.1459 times expert 0's output, 0.1249 times expert 1's, and so on. Rather like this: Diagram 4 : An MoE with all experts active With that, we have something where the router is on the backward pass through the computation graph from the output context vectors (and thus from the loss). The weights provide a route back, so the router will get trained. But that version, of course, is running all of the experts for every token, so it doesn't have the computational benefit of a modern sparse MoE. It also has an issue, as they point out in the paper, that if you imagine that you want each expert to have a well-defined responsibility, things get messy. Imagine that expert 2 in the above example was the one that knew how to solve a particular task; all of the other experts would be contributing too, and unless you find some way to train their weights down to exactly zero for that task, to get a good loss on your training run they'd need to learn to balance each other out. Their solution was to run all of the experts then to use a gating function on the output -- that is, it would drop all of the results from the non-selected experts (see their figure 1). But the modern solution is a bit different; for that, let's move on to "Outrageously Large Neural Networks". What we really want is an output from the router where some of the weights are zero. Specifically, with our k active experts, n total, we want all but the top k of them to have a weight of zero. Then when we're running the network, we can skip those zero-weighted experts and get a result that will be identical to what we would have got without skipping them. "Outrageously Large Neural Networks" does this with a trick that will be familiar from the causal mask in the GPT-2 code. Let's call the "raw" weights from our router : Let's run that through softmax again: That gives us some initial weights. But we want all but the top k to be zero. Now, if we want something to be zero after softmax, then it needs to be − ∞ on the way in. I'll show the code to do this later, but for now, let's just assume that we have some magic to set all but the top- k values to that. In our concrete example with k = 2 , the original with that change applied would be: Now we can run that through softmax: And we have some weights! With those, we can conceptually run this "all-experts-active" kind of network: Diagram 4 (repeated): An MoE with all experts active ...but skip the experts whose weights are zero. The weights that we provide through the steps above will mean that the router will get trained. That's pretty neat! There is one thing to highlight, though. Imagine if we have one active expert -- that is, k = 1 . We'd mask out all of the other ones: ...and softmax: The weight will always be one, regardless of the original logits. And if a function can only ever return the same result, its derivative will always be zero, so there will be no gradients and it can't be trained. Interestingly, that's something that is -- almost silently -- addressed in the Switch Transformers paper. They are specifically looking at k = 1 MoEs, and they mention the "Outrageously Large Neural Networks" paper's routing system, then present their own calculations which instead of replacing non-top- k values with − ∞ , and then softmax, do the softmax first, then zero out the non-top- k . They don't highlight the difference. I decided that for this experiment, I'd go ahead with the non-Switch Transformers calculations, anyway, and do the replace-with- − ∞ -then-softmax system, with some guard code to prevent it from accidentally being set to k = 1 . Most recent MoE models I've seen have at least two active experts, after all. The good news is that not only did it work -- when I checked later, it also matched the choice taken in Mixtral, which feels like a solid endorsement. 2 3 So, at this point, we know how an MoE works in theory. Let's start coding. I started with the code that I had from " Build a Large Language Model (from Scratch) ". In that, the Transformers block looked like this: I wanted to keep the capability to run this in MoE or non-MoE mode, and decided that I'd do that by extending the model config in so that it would have an optional section, which would include MoE-specific stuff. If that wasn't present, then I'd create a normal dense LLM. To do that, I replaced the line in that assigned to with this: Next, it was time to implement the class itself. The obvious config that it would need was how many experts there were in total, and how many were active per token: Now, there's that issue where having just one active expert will pin the softmaxed weight to one, so it won't train -- so I decided to protect against that (and against another obvious mistake that the config could contain): Next, it was time to create our router, mapping from the incoming context vectors (of size in this code) to the number of experts: I decided to make it unbiased because I have a vague impression that that is the fashion these days -- nothing more principled than that :-) Next, we needed the experts themselves, each of which would be one of the same modules as we were using in non-MoE mode: They needed to go into an because if they were just in a regular list (as I discovered when I tried it) they would not be registered by PyTorch as things containing parameters belonging to the module. We create our optimiser for training with code like this: ...and so anything that isn't in will never get updated, which would be a Bad Thing. That was enough to have the pieces in place. It was time to write the forward pass -- to get the router logits, do the top- k , softmax, and then to use those results to run the experts. Getting the logits was simple enough. Looking at the method: ...we have an incoming set of context vectors, . That is shaped . Let's try to visualise that. Understanding tensor operations is something you can kind of short-cut with intuition, but I think that in order to really understand the next steps it's best to have something a bit more tangible in mind. You can think of an order-3 tensor like this as a cuboid. With of 3, of 5, and of 7 -- artificially small values to keep things simple -- it might look like this: Diagram 5 : The tensor as a 3-D cuboid Every dot in that cuboid is a single number. A single context vector is the numbers as you follow a line from the "front" of the cube -- the × face to the left -- to the "back". That 3-D representation was fiddly to get right, and is likely to get confusing if we keep using it. So instead, let's look at two 2-D views of the same thing, from the front (where you are looking at a × face), and from one of the sides, where it's × : Diagram 6 : as two 2-D views So in this view, each of the circles in the "Front view" is the first number in a specific context vector, and each of the rows in the "Side view" is the set of numbers that make up a specific context vector. Hopefully that's easy to visualise. Now, a PyTorch layer like our operates on the last dimension in the tensor you pass in. For our , shaped , it will work on the dimension -- that is, it will operate on each context vector independently, which is what we want. That gives us a set of logits shaped . In our visualisation, that's a × × cuboid; let's diagram that with four available experts: : Diagram 7 : as two 2-D views Again, if we look at the front view, each circle represents the first element of the routing logits for one of the context vectors, and in the side view, each row represents all of the routing logits for a context vector. So we have the right data in our cuboid. From what we're calling the front view, it's got the same dimensions as , which is useful. In order to work out which of our experts each context vector should go through, we need to get the top- k values for each context vector in . PyTorch has a function to do exactly that: The tells it to work across the last dimension, which is the dimension of length . It returns the top values and their positions, as tensors -- both shaped . So, we have two new tensors, which we can visualise as cuboids, both of × × . Let's show that with : Diagram 8 : (or, equivalently, ) as two 2-D views Both and are shaped that way. As normal, the "front" face is the same; each circle corresponds to a context vector, and for it holds the highest value in that context vector's logits (because returns results sorted), while for it holds the index that that highest value sits at in the logits list. Looking at the side view, we're seeing a list of top-k logit values or indices for a given context vector. Now, we want a version of that we can run through softmax to get some weights, and in order to do that we need to replace the non-top- k values with − ∞ for each of the context vectors. The solution that I hit on eventually (after various other attempts 4 ) was this. Let's start with a tensor identical in size to , but full of − ∞ s: Now, contains the values that we want to have in there, and contains the indices in the logits lists where they should go. All of the other values can remain as − ∞ . This will, of course, be the same shape as : Diagram 9 : as two 2-D views Both and are compatible in their first two dimensions with , and thus with -- that is, their front faces in our visualisations are the same -- compare the front face of the above with the one for diagram 8 . For all of our tensors, the first two dimensions correspond ultimately to an incoming context vector in . It's just the last dimension and the data they contain that differ. PyTorch has a function on its class, which takes a list of positions and list of values. It takes one specified dimension, and treats that one essentially as a set of lists. Then it takes two other tensors with the same number of dimensions as each other, each of which matches in size in all but the specified dimension, and treats the other dimension as being a list of values for one parameter, and list of indices for the other. It overwrites the data at the specified indexes with the specified values. That sounds useful! Let's make it concrete. Remember that in all of these cuboids we're visualising, the front view we're looking at corresponds ultimately to a context vector. In , it's the raw logits for expert routing for that vector -- the number on the front face is the first of those (corresponding to the logits for routing to the first expert). If we look at the side view, each row would correspond to the set of logits for a given context vector. Now, let's consider just that. Inside , our specific context vector might have logit values corresponding to it like this: Diagram 10 : A single CV's routing logits in-place in the cuboid Visualised the same way, with k = 2 the equivalent part of would look like this: Diagram 11 : A single CV's top-k routing logits indices in-place in the cuboid That is, the function identified that index 2 in the original logits was the highest value, and 0 was the second-highest. Likewise, would have this: Diagram 12 : A single CV's top-k routing logits values in-place in the cuboid These diagrams are getting a bit unwieldy; let's look at this specific context vector's data as lists. For our selected context vector, we have these logits from diagram 10 : ...these top- k indices from diagram 11 : ...and these values from diagram 12 : Our tensor is just full of − ∞ s, and is the same shape as , so its corresponding part is this: What will do is select indices 2 and 0 (because of the values in ), and copy the corresponding numbers from on top of whatever is already there: It will do that for every one of the positions in that front view. So now we have what we wanted: a grid of routing logits for each context vector, where the non-top- k ones have been replaced by − ∞ . That's pretty nifty! And so here's the code to do it: I was initially a little worried that the whole "replace the logits with a 'static' tensor and then scatter the values in there" approach might break the computation graph, but tests showed that it didn't. My intuition is that because the − ∞ s were summoned out of nowhere, they are dead ends for backprop, but because the logits that we're copying in come from previous calculations, they are not. So once we've done that, we have our top- k logits, ready for a softmax -- so it's time to do that: ...and we have our weights, yet another cuboid like this: Diagram 13 : as two 2-D views Just as before, a given number on our front face is the first of the expert weights for a specific context vector, and each row in the side view is all of the expert weights for that context vector across all experts. Post-softmax, the weights for active experts for that context vector are positive numbers, and the weights for the inactive ones are zero. The next step was to actually feed the context vectors into their experts. I decided to keep this simple, and just iterate over the experts, one by one, and for each one to find which incoming context vectors wanted to go to it, run them all through, and then reassemble the results. I believe that larger MoE systems have smarter routing systems -- for example, for a huge model that can't fit on a single GPU you might have different experts on different GPUs or even machines, and route to them in parallel. But for my toy-sized models, this felt simplest, and felt like it would be efficient enough 5 . The way I decided to do this was to use an "accumulator" model. We know that the shape of the outputs is the same as the shape of the inputs -- that is, if we have our incoming shaped , then the output will have the same shape. So I started off by creating a tensor of zeros of that shape: So, just like in diagram 6 , it will look like this: Diagram 14 : as two 2-D views ...but it would be filled with zeros in every position. The plan was that each expert would be run on the appropriate context vectors, its outputs would be scaled by the weights that were calculated by the router and its associated top- k and softmax for that context vector/expert pair, and then the results could be added into . That would give us the result we wanted: after all experts had been run on their respective context vectors, would have the weighted sums. So the next step was to iterate over the experts: Now, we need to know which context vectors wanted to be fed to this expert. Let's take a look at our representation of again: Diagram 13 (repeated): as two 2-D views We can see it as a bunch of "slices", each the same shape as that front face, , where each one is the weights for a given expert. The front face itself (corresponding to the rightmost column on the side view) is the weights for expert 0, the next one "back" is for expert 1, and so on. And in code, we can get the slice for the expert with index like this: That will be a simple 2-D grid of numbers like this: Diagram 15 : sliced for one expert You can see that it's a grid of one number for each context vector in our input -- the same shape as the front face in all of the diagrams so far. Each number is the weight that the expert with index has for the context vector in question. We can now do a comparison: That will give us a new grid of the same shape, but the numbers have been replaced with booleans -- if the corresponding context vector has a weight for this expert that is greater than zero, otherwise. We've got a mask that identifies exactly the context vectors that we want to run through this expert. Let's imagine it looks like this for some particular set of context vectors and some specific expert; I've coloured in the circles representing and left the ones white. Diagram 16 : what might look like for one expert Now comes the clever part :-) Remember that our original incoming context vectors, in , looked like this: Diagram 6 (repeated): as two 2-D views We can use our mask -- the grid of s and s in diagram 16 -- to select a subset of the context vectors in there, like this: That will return us the subset of the incoming context vectors that we want to run through this expert. In terms of our diagrams above, it will be selecting the context vectors in the front face that are "selected" by our mask, and taking the "cores" as it goes back through the cuboid from there. The question is, what shape will it be? You can see that there's no simple 3-D shape it could be. The first sequence in our batch -- the first row on the front face -- has two selected context vectors, while the second has one, and the third three. You could imagine a world where the output would be the same shape as the input -- that is, a tensor the same shape as -- but with the non-selected numbers replaced with s or something like that. But instead, PyTorch produces what amounts to a list of the selected context vectors for this expert. We'll call the number of selected vectors , and it's six in the example diagram above, so the shape will be , like this: Diagram 17 : the "selected" context vectors Now, note that this is a big change from all of our tensors so far. It has lost the connection back to the original tensor's shape. In all of the other tensors, we could map something back to the original context vectors in . But with this new one, we have a bunch of context vectors with no inherent connection to where they originally came from in . We'll come back to that. But for now, we have the data that we want to run through the expert with index . The expert is an FFN -- our two linear layers with a GELU in between them -- and will treat all but the last dimension in whatever we feed it as "batch" dimensions, so it will work over that dimension, as we want it to. So the code to actually select the context vectors -- our above -- and then to run it through the expert itself just becomes this: We get the results of running our selected context vectors through the expert, still shaped . (If you're familiar with the GPT-2 code, that in the code might seem odd. We'll come back to why the expert is returning a tuple and why we are ignoring the second item in it later.) So we have our results -- we've run the appropriate context vectors for this expert through it. But they're in a slightly funny shape, and we're going to need to fix that later on. But first, we need to apply the appropriate weights for this expert to each result. Remember that is a grid of booleans, , like this: Diagram 16 (repeated): what might look like for one expert ...where -- filled in the diagram -- means that this expert is active for the corresponding context vector, and means it isn't. looks like this: Diagram 13 (repeated): as two 2-D views Now, previously we did this: ...to pluck out the context vectors from where the mask was true. If we were to do ...then we would be doing something similar. We'd get a grid like the one of context vectors in diagram 16 , with one row for every selected context vector for this expert, except that instead of each row containing a context vector, it would contain that context vector's per-expert weights: Diagram 18 : all expert weights for the "selected" context vectors Now, on its own, that's not particularly useful -- but you can hopefully see that column is the weights for our expert for all of the context vectors that are going to be run through it! So if we do the same masked lookup as before, but also select that column, we get code like this: What we're saying is "pick each item in that matches a in , then take the th element of it". It will look like this: Diagram 19 : this expert's weights for the "selected" context vectors That is exactly what we need to get the weights for our results! We need to multiply the i th element in -- an output context vector that is the results for a particular input context vector -- with the i th element of . However, there is one tensor-compatibility issue. What we have is shaped -- that is, it's essentially just a list of numbers. Now, we want to multiply our results by these weights, which means that we want to broadcast them across , which is shaped . Naively you might think that if you try to multiply a tensor by a one, PyTorch would match up the two dimensions of the same size and it would just work. But it would actually try to match up across dimensions from right to left, and would complain that does not equal . So instead, it's best to feed it an explicit tensor so that it knows which dimensions we're trying to match up, and which ones it should broadcast over: ...and that will give us of shape That would look identical to diagram 19 in the way I've been diagramming these things, but it's technically different and necessary. So now, with the having fixed the dimensionality so that we can do a broadcast, we can do this: ...and with that we have our weighted results for this expert -- all of the context vectors that should have been run through it have been, and they've been multiplied by the weights that the router gave for them, so we have our backprop channel for training the router. The next step is that we need to somehow put these results into . Specifically, because it's initially all zeros, and we want it to hold the results of adding together the results from each active expert for each context vector, we need to add them to whatever is already there at this point in the iteration through the experts. Even though -- as noted earlier -- we have, by this point in the calculations, lost the connection between the tensors we've been working with and the original "front-facing" grid of our original tensors like , where we could link something directly to the incoming context vector that it related to, it's actually surprisingly easy to patch that up :-) Remember that outside our loop through the experts, we created like this: So, just like in diagram 6 , it looked like this: Diagram 14 (repeated): as two 2-D views Now, previously we had code to extract the context vectors that we wanted to go through our expert from , and it looked like this: That gave us what amounted to a list of context vectors, like this: Diagram 17 (repeated): the "selected" context vectors Now, the interesting thing about a masked lookup into a tensor like is that not only can you do it to extract a subset of the values from the tensor like we did with -- you can also use it to assign to the masked subset of the elements in the tensor. To make that concrete: when we did We got a tensor sized . That means that if we were to do: ...then given that the mask is the same, and has the same shape as , then we'd also get a result. But the thing with assignment means that if we had some tensor shaped -- let's call it -- then we could do this: With that, instead of selecting the parts of and using them in the future, we're overwriting them with whatever is in . Furthermore, we can use the same trick with the augmented assignment operators like : ...means "select the elements from that have in , and then increment them by the elements of ". So finally, we get our code: That multiplies the results from the expert by the corresponding weights, and then adds them in to the running totals we're keeping in . Because is the same shape as , and thus is the same shape as , and we've kept that same shape as we went through the expert itself and multiplied by the weights, it will work. And with that, we've done all of the calculations that we need for this specific expert, so we can go back round the loop for the next one. When we've finished with all of the experts, we can return the result: Phew! That was quite a lot of explanation, but I think that what is going on in the code needs it. I originally wrote it in a kind of flow state of inspiration, and was very pleased with it, but it feels like now that I've explained it, I actually understand what my subconscious must have known while I was writing it. And I hope it's reasonably clear for anyone reading this. But there was one extra thing I wanted to keep track of before I started running this: the balance between the different experts. We'll come back to this in some detail later, but a problem with MoEs is that they can wind up depending heavily on specific experts, and ignoring the others -- the auxiliary loss I mentioned way back in the intro to this post is required to avoid that. I didn't want to implement that yet, but I wanted to log enough data to see if it really would be necessary with my setup. A good way to keep track of how much each expert is being used is to record the logits -- the original results from the router, before the top- k and the softmax -- and the actual post-top- k , post softmax expert weights that we actually used. The changes to do this are dotted around a bit, so I've linked to the appropriate lines on GitHub and if you have the screen real-estate to do so, I'd recommend that you use that to follow along. However, I've tried to put enough code inline in this post below that it should be comprehensible without that. I decided that I'd return the logits and the expert weights from the module's forward pass as an extra output here : Now, back in the we were setting to either a object or a depending on whether MoE was enabled for this model here : That meant that in our forward pass where we previously just did this (I won't link to this because it's an old version and having links to different versions would be confusing): ...then if the model was an MoE, we'd be getting a tuple -- the actual outputs from the module and then that extra routing information we were returning. But if it was a , we'd just get the outputs. I decided that the simplest fix was to change the module so that it returned values that were compatible with . The old , which was this: ...became this : That meant that we could do this in place of the code above: If we had an MoE in , then we'd get some real MoE routing info in , whereas if we were running a normal dense model then we'd get . This, by the way, explains the odd bit of code in the for that you might remember from earlier -- this bit : there is one of the modules, so what we're doing there with the underscore is just ignoring the that we know we're returning from it. The next step was to work out how to get the extra information that we had in out of the , if was a . Now, is used in an in the class here : That means that the outputs from the method of the first one are fed directly in as the inputs to the second one, and so on. Additionally, assumes that the forward method just takes a single input. What I really wanted at the end of this was a list of pairs, one for each Transformers layer. So I decided to pass an accumulating list into (allowing it to be ), like this : ...and then append whatever routing info came from the (potentially ) FFN to that here : ...and then return it from for the next layer here : Then, in I wanted to return it to whatever called the model for inference. I didn't want to change the implicit API that I was using -- pass in inputs, get next-token logits -- so I decided to just attach it to the like this : I'm not sure, in retrospect, that "smuggling" the routing info out like that was the right choice. Perhaps a Hugging Face-like model where the LLM returns some kind of "output" class that contains and other stuff like this as explicit fields would be better. But this setup seemed to work for my use case, so I've left it as-is for now, with a mental note to revisit later. Next, I modified my training script to store the that we got when we did our forward pass in a list, and then to store that in the metadata associated with my checkpoints for later analysis. The training script I'm using is something I developed after working through Raschka's book, and it's kind of complicated -- although the core is essentially the same as the one we use to train our model in chapter 5, I've built it out to allow training across multiple GPUs , with all kinds of tweaks that I've learned about since. So I won't dig into that code in any depth in this post; if you've been following along with my various training posts in the past, though, and want to see how this all fits in, you can see the code that accumulates the routing information here (note that it uses so that we don't accumulate compute graph information over time), the code that passes those details into the checkpointing function here , and the updated checkpointing function here . So: with those changes, I had a plausible-looking MoE setup. I was confident that it would be able to train, but suspected that it would not be able to balance load between the experts well and would collapse to using a subset of them. It was time to give it a go. The first question was, what total number of experts, and how many active ones, should I have? I was pretty limited by what I could fit into my RTX 3090's VRAM, and what an appropriate training speed might be, but after a bit of fiddling around I came to the conclusion that six experts with two active would fit into memory and train reasonably quickly, and that didn't sound like a crazy balance. Mixtral is 8 experts, two active, for example. I wouldn't be able to fit in the batch size of 6 that I had used in the past when training 163M-parameter dense models; the largest batch size I could fit in was 3. (For those who've been following my previous LLM training, because I'm using gradient accumulation over 16 steps with the microbatch size of 6 -- for a global batch size of 96 -- I could get the same effect, and fit the MoE into VRAM, by going down to a microbatch size of 3, with 32 gradient accumulation steps.) The next question was how many tokens to train for. As an experiment I kicked it off with the same 3.2 billion tokens that I would normally use to train my 163M-parameter models. I knew that this larger model would need to be trained on more tokens than that, but I just wanted a reasonably serious run to see how the model behaved. The training script predicted that it would take just over two days to complete this experimental run, which was perfect. I had reached this point during the late afternoon on a Friday, and had stuff to do over the weekend that would keep me away from my computer, so that run length would mean that it would be ready for me on Monday. On Monday, I came back to see that it had completed properly: Over the training run, the loss declined nicely: Out of interest, I ran an evaluation script to see how it performed on the held-back test set that I normally use to compare my models: That ranked it better than the 163M-parameter models I'd trained using PyTorch, but worse than the original GPT-2 small, and also worse than the models I've trained with JAX (which I believe got lucky with their initial untrained random weights). That was a pretty solid result, but as this was an exploratory training run, there's no point in putting much weight on it. For a start, this model was clearly undertrained. The number of tokens, 3.2B, was Chinchilla-optimal for my 163M-parameter models, but this one had more than twice the total parameters at 446M, and had 220M active for each token. Larger models need more tokens to be well-trained. The more important result was the load-balancing between experts. I put together a notebook to ingest the metrics that I was writing to the checkpoints, and to create two charts for each layer: In the charts, for each global step, I plotted a line for each expert; if it had an average probability less than uniform / 3 -- that is, it was getting less than a third of the tokens that a perfectly flat distribution across all experts would give it -- then the line was white, meaning that it was being "starved". If it was getting a number of tokens between uniform / 3 and uniform * 2 -- then the line was green, because it was getting what felt like a reasonable number of tokens. And finally, if it was higher than uniform * 2, it was plotted in red: it was being overfed. The charts showed exactly the problem I was expecting to see. Let's look at layer 0: You can see that it started off pretty well-balanced, but rapidly came to prioritise expert 5; the remainder of the probability distribution looks like it must have been scattered over the other experts, with a slight preference for expert 3. The other layers came back with similar issues: certain experts were strongly preferred while others were starved: So the problem was real. The model was relying heavily on some experts and ignoring others. We needed to do extra stuff to balance the load across the experts. The problem with load-balancing across experts has been clear for quite some time, and it was covered in "Outrageously Large Neural Networks" back in 2017. The solution they used is simple, and clever: define an "auxiliary" loss, which goes up as the experts become more unbalanced, and down as they become more balanced. You can then add that (scaled by some amount) on to the normal cross entropy loss that you get by comparing the outputs from your model to the training targets, to get a combined loss number -- and then you back-propagate using that combined loss. By reducing the combined loss, then you optimise both for correctness -- is the model doing its job as an LLM -- and for balance across the experts. Of course, you need to be careful about the scaling factor that you use to adjust the auxiliary loss before adding it. If it's too small, then the load-balancing will be de-prioritised and you can wind up with imbalanced experts anyway. But if it's too large, the training process will prioritise balance over quality of the output, and you'll wind up with a crappy model that balances load across the experts beautifully. We'll come back to that shortly. The "Outrageously Large Neural Networks" paper's particular calculations for the auxiliary loss are a little complicated -- they generate two separate numbers and combine them -- and it was simplified in the GShard paper, and then again in Switch Transformers. I found the last of those reasonably easy to understand, and decided to adapt it. Let's start off with their formalisation -- but keep in mind that they had only one active expert per context vector, so we'll need to make some changes. Firstly, they define a per-expert number, f i -- the i subscript means that it's for expert i . That's a little intimidating-looking but is much simpler than it looks. They define it as "the fraction of tokens dispatched to expert i ", and in that light we can interpret it. The x ∈ ℬ that we're iterating over in the ∑ is basically: So, the stuff inside the ∑ is done once for each of our input context vectors across all sequences in a batch. The 1 { . . . } (which should be rendering with a kind of doubled-1 -- Chrome, as of this writing, doesn't handle that, though Firefox does) is meant to mean "1 if the condition in the braces is true, 0 if it's false". In argmax  p ( x ) = i , the function p is the "raw" probabilities from our router. As I mentioned earlier, they were doing softmax and then zeroing out the non-top- k values, so they had those raw numbers available. We'll come back to that in a moment. But for now, the condition inside the 1 { . . . } just means "true if this expert was the top one for this particular incoming context vector, false otherwise". Putting that all together, then we're just counting the number of tokens in our batch for which expert i is the top pick of the routing network. The 1 / T at the start then divides that by the number of tokens in the batch (which is what they mean by T ), and for a network with one active expert like the Switch Transformers ones, we've got -- as they say -- "the fraction of tokens [in this batch] dispatched to expert i ". So that works nicely in the one-active Switch Transformers world. But we can fairly simply extend it to handle multiple active experts, while keeping similar semantics. Let's say that we calculate how many of the context vectors in a batch were sent to expert i ; we can divide that by the number of tokens to get something equivalent. It still means "the fraction of tokens dispatched to expert i ". (Ab)using their notation, if we say that our number of active experts is k , then it might look something like this: Let's run with that for now. As well as f i , they define another per-expert number, P i : They describe this as "the fraction of the router probability allocated for expert i ". Again, we're iterating over every context vector in every sequence in the batch -- but here we're just adding together all of the raw probabilities for the expert for each one, then dividing by the number of elements in the batch. In other words, we're working out the expert's average probability across the entire batch. Much easier! And in our multiple-active-expert world, their equation works -- it means exactly the same thing. Now, both of these calculations, as they're expressed in the maths, depend on something that we're not currently calculating. Remember, Switch Transformers worked out the weights for each expert by doing a softmax across all of the raw logits that came out of the router's linear layer, then zeroing out the ones that were not in the top- k -- that is, for their k = 1 setup, all but the highest (argmax) one. So they had those "raw" softmaxed probabilities knocking around. But because we are replacing the non-top- k logits with − ∞ and then doing softmax, our router weights aren't the same -- they're just the relative probabilities of our selected k experts. For f i that doesn't matter; we can use our weights and easily identify the experts that were routed to for a given element in the batch, because they're the only ones that are not zero. But for P i we do need the pre-top- k 's logits from the router. And, not entirely coincidentally, we already have the code to make that available! In order to do those charts above -- the ones showing where experts were being starved and when they were being over-fed in the non-load-balanced training run -- we passed the raw, non-top- k 'ed, non softmaxed router logits, and the expert weights after the top- k and the softmax, out of the module's : ...and then added code to feed that through to the training loop because it was needed for the metrics that we used to generate those charts. The numbers that we kept for the "raw" side of things were the logits rather than the actual probabilities, but we can fix that with a simple softmax. So that means that in our training code, we already had the numbers to work out P i and f i for our experts. They needed to be combined to make up a single scalar auxiliary loss for the model when run over a batch, and Switch Transformers does that like this: If you imagine f as a vector containing all of the per-expert f i s, and P as a similar one containing the P i s, then that ∑ is a simple dot product, f · P . They then scale that up by the number of experts N , multiply by a scaling factor -- the one I mentioned earlier to balance load-balancing auxiliary loss against "real" cross entropy loss -- which they call α . (We'll come back to α later.) So, we have a set of calculations to work out an auxiliary loss; there are a few extra things I'd like to highlight before we dive in to the code. Let's imagine what a "perfect" router would look like from the perspective of these calculations, firstly for the Switch-style one-active-expert model, and then for our own more general case. Starting with f i ; it's the number of tokens in the batch for which expert i is the chosen one, divided by the number of tokens in the batch. We want all of our experts to get the same amount of "traffic", so for N experts, one active per token, clearly each one will be active 1 / N of the time. So that should be the value of f i if everything is balanced. Now let's think about P i . Again, we want each expert to receive 1 / N of the tokens -- so its average probability should also be 1 / N . So, that means that for perfect balance, all of our P i s and all of our f i s should be 1 / N . That means that when we do the ∑ in that loss calculation to work out the dot product, then for a perfectly balanced router, each one will contribute: There will be N of them, so that will come to ...and then we're multiplying by N to get the loss, so the result (disregarding the scaling factor α ) will be 1 . So, when balance across the experts is perfect, the Switch Transformers auxiliary loss has a value of 1 . However, they have only one active expert, and our equation is slightly different to allow for the fact that we have k of them. I won't go through the boring derivation again, but if we replay the maths above with k active experts, we get an auxiliary loss for the ideal, perfectly balanced router of k . That's not a problem in and of itself -- after all, this is just a number we're trying to minimise, and it's not super-important what we're trying to minimise it to. But it does matter when we're talking about α , because if the auxiliary loss is larger, we'll need to scale it down more to stop it from "drowning out" the signal from the actual training loss. The Switch Transformers paper explains what values they used for the scaling, but ours will be different because of the different number of active experts. The auxiliary loss calculation above only covers what happens in one layer, so we need to work out how to combine the contributions from all of the layers. In the paper, the only mention they make of multiple layers in this context is: For each Switch layer, this auxiliary loss is added to the total model loss during training I took that to mean that we just sum up the scaled auxiliary loss across all layers and then add that sum to the normal cross entropy loss. I think that's the most natural interpretation of what they are saying. So -- given the value of k for a perfectly-balanced router that I worked out above -- for the model I was planning to train, with 2 active experts, 12 layers, the auxiliary loss before scaling by α would be 24. This, by the way, is where my implementation differs from Mixtral's -- or, at least, the Hugging Face source code for it as of this writing. In that, they do something that feels a bit odd. They treat (for example) expert 1 on layer 1 as being the same as expert 1 on layer 2, and so on throughout the layers, then do the calculations just once. That feels a bit dodgy. After all, you can imagine that expert 1 on layer 1 might be being starved but its equivalent on layer 2 might be getting overfed, and the two would balance out. I don't know if that's an error in that implementation, or if there's something I'm missing. Conceivably I might test it some day by training another model using their loss function, but I suspect I won't get around to that. There's also something that made me hesitate a bit in the Switch Transformers paper, and which I think I still need to ponder a bit. That calculation for f i : ...did not look differentiable to me. Things like argmax (and our own equivalent's is-in-top-k ) are generally not. Indeed, they confirmed that shortly after defining it, but in a way that gave me pause: The objective can also be differentiated as the P -vector is differentiable, but the f -vector is not. The "objective" they're referring to is the auxiliary loss, and it makes me a bit uncomfortable that something that is defined in terms of A and B is differentiable if A is, but B isn't. My hand-wavy way of thinking about it right now is that the undifferentiable bit can be treated as a constant, so as long as part of the calculation is differentiable, the whole thing can be treated as such. But I'm not 100% happy with that, and need to think further. But now, I think, we've covered the maths for the auxiliary loss, so it's time to dive into the code! My old training loop had the following code to do the forward then the backward pass: Let's strip out all of the extra enhancements that I have accumulated there on top of the simple training code from the book; there's AMP , DDP and gradient accumulation and if we remove that it would simply look like this: Hopefully that's familiar! What I wanted to do was add in the auxiliary loss (if we were training an MoE), scaled by that α scaling factor. What I came up with (and again, here I've stripped out all of the stuff required by those enhancements): Note that I wanted to keep track of the router losses in that list as well in order to monitor them as the training run progressed, just like I normally monitor training loss (as you can see in the loss chart above). That's all pretty nice and simple -- if we are getting MoE routing info back from the model, then we call this new to work out the loss from the maths in the last section, and then add it on, scaled by , which is what I decided to call the somewhat-opaquely-named α from the Switch Transformers paper. Adding all of the AMP, DDP and gradient accumulation gubbins back in, the final code looked like this : So that was simple enough (for LLM-training values of simple). The code to route the from the training configuration file is not really worth going through, and nor is the code to save average, minimum and maximum values from into the checkpoint metadata, or to chart those (though I'll show the charts later). The interesting bit is, of course, that function. It's here and looks like this: Let's look at it from the outside in. We start with a list called . Remember, this has been passed back from the MoE model for a single forward pass of a batch. It contains one item for each Transformers layer in the model, and those items are pairs of . We start off with a total routing loss of zero, and then for each layer, we do some stuff to work out its routing loss, and then at the end of the loop, we add it on to our running total. Finally, we return the total. Obviously, the fun stuff is inside the loop :-) It's time for some of those tensor diagrams again. Let's remind ourselves of the shape of : Diagram 13 (repeated): as two 2-D views Each cell on the front face relates to one context vector in one sequence in our batch, and the "core" going into the cuboid from there (horizontally along the side view) is the weights we actually used when routing the context vector in question to the experts -- zero for unselected experts, some value between zero and one for the selected ones. Now let's go back to the code for a moment. We start off by using the shape of to work out what our different dimensions are: Then we do our first block of calculations, trying to work out f i from the maths, "the fraction of tokens dispatched to expert i " We want to do this efficiently with a vector calculation, working out all of the f i s (remember, there's one for each expert) in parallel. Our first step is to flatten out the grid into a single dimension -- essentially stacking all of the front view's columns on top of each other to make just one long column, like this: Diagram 20 : flattened, as two 2-D views ...or, more simply, as it's now just a 2-D Tensor, like this: Diagram 21 : flattened, as one 2-D view We can use PyTorch's method to do that without having to copy any data around in memory -- as the name suggests, it just returns a different view on the same data: Now, remember that these are the expert weights. Each one of those cells contains a number -- zero if the expert was not selected for the context vector that corresponded to it in the original layout, or some non-zero number if it was. So if we do this: ...then we'll get a tensor of the same shape, where we have if the weight was more than zero (that is, the expert was active for that context vector), otherwise. And that means that if we sum down those columns in diagram 21 , we'll get a new grid of one row, and columns, which represents the total number of times each expert was active in the given batch: Diagram 22 : as one 2-D view So, in code, we can just do this: If we divide that by the number of context vectors in the batch, we've got a new tensor, a row with columns -- the same shape as diagram 22 -- containing exactly what we want: That is, is a vector f in terms of the maths, containing all of the f i s that we want for our auxiliary loss calculation -- that is, all of the result across all experts for this: The calculations for the P i s are very similar. We start off with like this: Diagram 7 (repeated): as two 2-D views So, each cell on the front face relates to one context vector, and the "core" going into the cuboid from there is the set of raw routing logits for that context vector, one number per expert. Firstly we need to convert the into probabilities by running them through softmax, along the last dimension -- the one that is the horizontal axis on the side view: Then we do the same trick with to convert the result to a single column of lists of length (metaphorically) -- the same shape as in diagram 21 : Now if we add them up across the rows like this: Then we get a single row, column result that has, for each expert, the sum of all of the probabilities it had across all of the context vectors in the batch, just like the one we had in diagram 22 . We can divide that by the number of context vectors in the batch: ...and that's our vector P containing all of the P i s, where each is one of these: Finally, we can multiply all of the f i s and their corresponding P i s by each other, and sum the results, by using a dot product, and then multiply the result by the number of experts: ...and that's this bit done (apart from the α ): And that's our auxiliary code wrapped up! We've been through , and we've already seen the code that scaled it by α aka , so we have a training script with MoE auxiliary loss using the maths in the last section! Again, I hope that the diagrams helped with that workthrough. The code is the kind of thing where it's easy to scan through and get a vague understanding, but I think that visualising what's going on step-by-step is important if you want it to really stick. With the code in place, the next thing to do was to explore what the right value might be for α . In the "Switch Transformers" paper, they say: Finally, a hyper-parameter α is a multiplicative coefficient for these auxiliary losses; throughout this work we use an α = 10 − 2 which was sufficiently large to ensure load balancing while small enough to not to overwhelm the primary cross-entropy objective. We swept hyper-parameter ranges of α from 10 − 1 to 10 − 5 in powers of 10 and found 10 − 2 balanced load quickly without interfering with training loss. But, as I noted earlier, that worked for them with their single active experts, but because I had multiple, my auxiliary loss would be larger. Now, given that their "ideal" per-layer loss was 1, and mine was 2 for my planned training run, it sounded like using half of their recommended value, 5 × 10 − 3 , would be appropriate. However, I was also a little concerned about the number of layers interfering with things. The auxiliary loss, as we saw above, was a single value per layer, all of which were added together. With my "ideal" per-layer loss of 2, and 12 layers, that meant that the ideal across all layers was 24. Now, they mentioned a specific value for α , but didn't mention the number of layers they had, or if they swept for different possibilities across different numbers of layers. That seemed strange! I decided to work on the hypothesis that they had found that as the number of layers increased, you needed the total contribution of the auxiliary loss to scale up in proportion. That was only a guess based on trying to fit together the info in the paper, though, and could well be wrong. I was in enough doubt, though, that I felt it would be wise to do my own, minimal sweep over a few values for α . For each, I'd do a one-hour training run. At the end I'd check the training loss -- that is, the pure cross entropy loss saying how well the model was doing at its real purpose of modeling language -- and the auxiliary loss, showing how well-balanced its usage of experts was. I got these results: I also generated router usage maps like the ones I gave way back in this post for the two-day training run with no auxiliary loss. I won't put them all in there, but: So, on the basis of those results (and, of course, the fact that it fit well with the Switch Transformers paper's recommendation), I decided to use α = 0.005 . It was time to train this thing! I decided not to think too hard about the right number of tokens to train it on. The Chinchilla number, 20 times as many tokens as parameters, is a heuristic that works for dense models, but is not meant for MoEs. You'd intuitively think that the "ideal" number of tokens to train a model with 446,410,752 parameters, 219,697,152 active per token would be somewhere between the Chinchilla-optimal numbers for those two parameter counts. But overtraining isn't necessarily a bad thing, so long as it's on non-duplicated data (or less than four epochs over the same data ), and I had a 10B token dataset all set up from my previous experiment. So training it on what would be the Chinchilla-optimal number of tokens if it was just a 446,410,752-parameter dense model didn't sound like a bad idea, so long as I could do it in a reasonable amount of time. That meant 8,928,215,040 tokens, which I had in my normal training dataset without needing multiple epochs. A quick check -- running the training script with that number of tokens configured -- told me that it would take 90,823 global steps to complete, over about eight days. It looked like each checkpoint would take up 5.3GiB, and I had about 348 GiB free on my disk, so I calculated that I could checkpoint every 1,500 global steps -- that would work out as roughly once every three hours. Not great, but losing a maximum of three hours work in the case of a power outage (or cat jumping onto the PC's power button) is not the end of the world. So I set things up with this model configuration and this training configuration , and on my dedicated training box, , I kicked it off: Just less than eight days later, it completed: The loss chart looked like this: You can see that the normal cross entropy loss (just tagged as "loss" on the chart) decreases nice and smoothly from random (about 10.82 with the GPT-2 tokeniser) down to that final training loss of 3.211, with only a couple of tiny spikes. I've also plotted the auxiliary loss for the MoE routing, on the right-hand Y axis, and you can see that while near the start there were a few bumps (and the max values for single iterations spiked up from time to time), the average was generally pretty close to 24, our "ideal" number. Indeed, for the last checkpoint, the average over all iterations was an almost-perfect 24.0847. The notebook that I had to plot layer-by-layer expert starving/overfeeding came back with some lovely green plots showing nice even routing, too: A couple of issues near the start of the run, but for the last 75% of it, they're all a sea of green. Lovely. I ran my normal smoke test , based on Raschka's from the book: what do you get if you ask your model to complete "Every effort moves you" with 20 tokens, with a temperature of 1? Reasonably coherent! It was time to do some evals and comparisons. I ran my normal loss eval : on a held-back test set of 19,200 sequences of 1,024 tokens each, what was the model's average loss? Comparing this against other models that I've trained, and the OpenAI GPT-2 small and medium weights, we get this (it's in bold): Not too bad, though not amazing. It was better than any of my 163M-parameter models, and OpenAI's GPT-2 small. But it was a bit worse than (but close to) the 345M-parameter OpenAI GPT-2 medium, which has fewer parameters -- albeit more active ones per token. I decided to do a second eval. I have one that I call the IFT test -- fine-tune the model on an instruction-following dataset, until loss starts rising on its held-back eval dataset, then run a test set through to get answers to questions the model has not yet seen. I then bundle together the responses from a bunch of models and ask GPT 5.5 to compare them. It's an extension of Raschka's example in chapter 7 of the book, modified to make it easier to compare different models. You can see the scripts here and here , and there are more details here . I kicked off the first script, to fine-tune the model and get its responses, and then handed that plus a bunch of responses from other models to the LLM for it to compare them. The results came back like this (note this this is sorted by loss, the "IFT rank" column is how well it did comparitively in the eval): As you can see, the correlation between loss and performance on this eval is interestingly loose -- I have an ongoing series trying to work out why that might be . In particular, OpenAI's weights consistently outperform mine, and I'm determined to find out why. But it was reassuring, at least, that the new, big model came in at rank 3, beating all of my other ones, even if it still lost to that pesky 124M-parameter OpenAI GPT-2-small. So, there we have it: a GPT-2 small model converted to an MoE with 6 experts per layer, 2 active per token. Its loss on the test set is pretty much where you'd expect, and its IFT eval makes sense, modulo the OpenAI weights weirdness. What does that mean, and what should come next? In this post, I started with the GPT-2 code from " Build a Large Language Model (from Scratch) ", and my own training script (which was originally based on the training code from the book), added on mixture of experts support including the auxiliary loss calculations that you need to make it balance load across its experts properly. After eight days of training, we wound up with a decent, capable model. So that's all quite satisfying in an intellectual sense, and -- at least in terms of how well it did on the test loss -- it landed pretty much where you might expect. And one thing I'm sure of is that grinding through the calculations has been a great work-out for my skills with PyTorch and tensor operations. But the interesting thing about MoEs is that they -- in theory -- provide similar performance to dense models, at a lower cost in inference computing time. I think there are some interesting follow-up experiments I can do. OpenAI's weights tend to beat mine, so if I exclude them from comparisons (until I've worked out why), I could try to build a mental model for whether MoEs are a good way to spend my scarce computing resources when learning more about LLMs. I could: I'm sure there are other options, and I'd love to hear people's thoughts on what they might be. Anyway, I hope this post has been an interesting journey, and explained things well. Any feedback much appreciated -- in particular, on whether the diagrams helped. I only recently added D2 support to my static site generator, and it's entirely possible that I was overusing my new toy :-) So: thanks for reading, and as always, comments and questions are very welcome in the comments below. To protect against linkrot, I'll put proper references to these papers, because they're important background. In terms of the implementation, Mixtral is a bit odd. It does the Switch Transformers trick of doing the softmax, then zeroing out the non-top- k values -- but then it scales up the resulting routing weights by taking the sum, then dividing each value by that sum (which will be less than one). But the net effect of doing that is exactly the same as the "Outrageously Large Neural Networks" technique of replacing non-top- k with − ∞ .  ↩ One other thing: the masking out of all but the top- k experts does make me feel a little suspicious. It feels in a hand-waving kind of way a bit like "dead-end" code in router that we had originally, where we didn't use the outputs to weight the sum at the end. And it looks like there is a real mathematical concern there (even if it's not quite the same); in the "Outrageously Large Neural Networks" paper they say: While this form of sparsity creates some theoretically scary discontinuities in the output of gating function, we have not yet observed this to be a problem in practice My intuition for this is that because we have the weights in the flow of the calculations, back-propagation can still try to (say) reduce the weight for an expert that should not have been used for a particular forward pass. That is a bit problematic, though, because (due to the softmax) you would expect that reducing one expert's weight would force the others' to increase -- the results of a softmax have to sum to one. On consideration the Switch Transformers softmax-then-mask system feels like a better solution to me, because although the weights for the unselected experts are "invisible" to the backprop, having been masked out, gradients that increase a selected expert's weight will kind of automatically decrease those of all of the other, non-selected ones -- and vice versa. I need to ponder this a bit more, and perhaps try another training run with the alternative Switch-style setup and see how it compares. One random idea -- the softmax-first approach means that the weight of the selected expert(s) will be lower than it would be otherwise, so that reduces the amount of signal that the FFN adds to the context vectors. But then maybe the FFN would just get trained to emit larger outputs...? But anyway, enough waffling: let's put that aside for now, and for the rest of this post I'll describe the code as I wrote it, using the "Outrageously Large Neural Networks"/Mixtral model for the router.  ↩ Here's one that I tried, and which worked, but had a non-obvious bug. Let's imagine we have these logits for one context vector: For two active experts, we want to replace it with this: By default, sorts the values in decreasing order (and keeps the indices in an appropriate order to match). So in this case, we'd have the top-2 values looking like this: So if you do something like ...then you mask out all logits that are less than the lowest value in the top-k list with − ∞ . Neat! The problem with that was that there could be a tie. Imagine if, for one context vector, instead of the values above, we had these logits: We want to select the top 2 experts, and the smallest top-2 value would be 0.4254, of course. But if we just replace all values less than 0.4254 with − ∞ then we get this: We have three active experts rather than two! That's not good. While I didn't think that hitting this problem would be all that likely in practice, it was a definite error, and needed addressing -- hence the solution I settled on.  ↩ Imagine an extreme case, 256 experts but one active per token. In a given batch, you might expect tokens to be scattered pretty much randomly across experts, so you'd need to run ~all of them. If each expert only uses a small amount of your GPU's power to run through the small subset of the tokens that are allocated to it, and you're processing one expert at a time, you'll seriously underutilise the GPU. But I figured -- and confirmed when I finally got all of this running -- that with the sizes of model and numbers of active/total experts I was using, this wasn't a problem. Another issue that might come up in really large models is that too much stuff in a batch might wind up going through the same experts. Although we will later on add code to make sure that on average each expert receives roughly the same amount of context vectors, within a given batch things might be imbalanced. To see why that might be a problem, imagine an LLM that is so large that you have different GPUs -- or even different machines -- handling different experts. Something in the trillions of parameters scale, for example. If in your batch, expert i winds up doing most of the context vectors, then the GPU it's on is a bottleneck. The setup I'm describing in this post is sometimes called "token choice" routing. In a sense, each token has chosen -- or, rather, the router has chosen on its behalf -- which experts it "wants" to go to. "Expert choice" routing inverts that; you generate the same logits, but instead of choosing the top- k experts for each token, you choose the top- k tokens for each expert, so that you guarantee balance across the experts. Both of these are something for future investigations, I think :-)  ↩ " Adaptive Mixtures of Local Experts ", the 1991 paper that introduced the name (but used a model where all of the experts were run for each input -- they were thinking less in terms of efficiency, and more in terms of having different parts of a network specialise in different things). " Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer ", from 2017, which showed that you could use gating on MoE models to reduce the amount of computation by only running the "best" of the experts for each input. " GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding ", from 2020, which built on the above, applied it specifically to Transformers models, and simplified some things -- more on that later. " Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity ", from 2021, which did some nice additional simplification. One showing how well-balanced the logits coming out of the router were -- that is, for each global step, how close was the model to having every expert getting exactly the same amount of utilisation in terms of the router's raw numbers. One showing how well-balanced things were after the top-k. This was a belt-and-braces thing: you can imagine a router that always favours a particular pair of experts, but only by a small amount, so the raw routing balance might look reasonably good (say, the two favoured ones get a probability of 0.2 and the others get 0.15 each), but the favoured two experts would wind up getting all of the "traffic". With α = 0 , as you'd expect to see from the very high auxiliary loss, things started going red and white pretty quickly -- it was focusing on specific experts and ignoring others. It didn't do too badly on training loss, though. With α = 0.001 , one tenth of the Switch Transformers number (and one fifth of what you'd expect to be ideal for ours by converting it naively), the load-balancing charts weren't quite so bad, but there was a speckling of red. I couldn't say for sure that it might not have managed to keep things reasonably balanced in a longer training run. α = 0.005 , however, looked really good! It actually had the lowest (best) training loss out of any of these options, and the auxiliary loss was not too bad, only 1.2 above the ideal value of 24. By the time we got to the α = 0.01 that was used by Switch Transformers (with their ideal auxiliary loss values that were half of ours), although the auxiliary loss continued downwards, training loss started worsening. This was certainly in line with what I would expect to see if the training started prioritising balance over actually creating a good LLM. Train a small model (say, the 163M-parameter size I've been messing with so far) on the same amount of compute as I spent on the MoE. How would it perform? My guess is that it would be worse. Train a dense 446M-parameter model on the same amount of compute and see how it matches up. It would be undertrained (it would need more calculations per token, so if I matched the compute, I'd have to train it on fewer tokens). I don't know if it would be better or worse. Train a Chinchilla-optimal dense model on the same amount of compute, and see how that did. To protect against linkrot, I'll put proper references to these papers, because they're important background. Jacobs, R. A., Jordan, M. I., Nowlan, S. J., & Hinton, G. E. (1991). Adaptive mixtures of local experts. Neural Computation , 3(1), 79–87 Shazeer, N., Mirhoseini, A., Maziarz, K., Davis, A., Le, Q., Hinton, G., & Dean, J. (2017). Outrageously large neural networks: The sparsely-gated mixture-of-experts layer. In International Conference on Learning Representations (ICLR 2017) . Lepikhin, D., Lee, H., Xu, Y., Chen, D., Firat, O., Huang, Y., Krikun, M., Shazeer, N., & Chen, Z. (2021). GShard: Scaling giant models with conditional computation and automatic sharding. In International Conference on Learning Representations (ICLR 2021) . Fedus, W., Zoph, B., & Shazeer, N. (2022). Switch Transformers: Scaling to trillion parameter models with simple and efficient sparsity. Journal of Machine Learning Research, 23(120), 1–39 . In terms of the implementation, Mixtral is a bit odd. It does the Switch Transformers trick of doing the softmax, then zeroing out the non-top- k values -- but then it scales up the resulting routing weights by taking the sum, then dividing each value by that sum (which will be less than one). But the net effect of doing that is exactly the same as the "Outrageously Large Neural Networks" technique of replacing non-top- k with − ∞ .  ↩ One other thing: the masking out of all but the top- k experts does make me feel a little suspicious. It feels in a hand-waving kind of way a bit like "dead-end" code in router that we had originally, where we didn't use the outputs to weight the sum at the end. And it looks like there is a real mathematical concern there (even if it's not quite the same); in the "Outrageously Large Neural Networks" paper they say: While this form of sparsity creates some theoretically scary discontinuities in the output of gating function, we have not yet observed this to be a problem in practice My intuition for this is that because we have the weights in the flow of the calculations, back-propagation can still try to (say) reduce the weight for an expert that should not have been used for a particular forward pass. That is a bit problematic, though, because (due to the softmax) you would expect that reducing one expert's weight would force the others' to increase -- the results of a softmax have to sum to one. On consideration the Switch Transformers softmax-then-mask system feels like a better solution to me, because although the weights for the unselected experts are "invisible" to the backprop, having been masked out, gradients that increase a selected expert's weight will kind of automatically decrease those of all of the other, non-selected ones -- and vice versa. I need to ponder this a bit more, and perhaps try another training run with the alternative Switch-style setup and see how it compares. One random idea -- the softmax-first approach means that the weight of the selected expert(s) will be lower than it would be otherwise, so that reduces the amount of signal that the FFN adds to the context vectors. But then maybe the FFN would just get trained to emit larger outputs...? But anyway, enough waffling: let's put that aside for now, and for the rest of this post I'll describe the code as I wrote it, using the "Outrageously Large Neural Networks"/Mixtral model for the router.  ↩ Here's one that I tried, and which worked, but had a non-obvious bug. Let's imagine we have these logits for one context vector: For two active experts, we want to replace it with this: By default, sorts the values in decreasing order (and keeps the indices in an appropriate order to match). So in this case, we'd have the top-2 values looking like this: So if you do something like ...then you mask out all logits that are less than the lowest value in the top-k list with − ∞ . Neat! The problem with that was that there could be a tie. Imagine if, for one context vector, instead of the values above, we had these logits: We want to select the top 2 experts, and the smallest top-2 value would be 0.4254, of course. But if we just replace all values less than 0.4254 with − ∞ then we get this: We have three active experts rather than two! That's not good. While I didn't think that hitting this problem would be all that likely in practice, it was a definite error, and needed addressing -- hence the solution I settled on.  ↩ Imagine an extreme case, 256 experts but one active per token. In a given batch, you might expect tokens to be scattered pretty much randomly across experts, so you'd need to run ~all of them. If each expert only uses a small amount of your GPU's power to run through the small subset of the tokens that are allocated to it, and you're processing one expert at a time, you'll seriously underutilise the GPU. But I figured -- and confirmed when I finally got all of this running -- that with the sizes of model and numbers of active/total experts I was using, this wasn't a problem. Another issue that might come up in really large models is that too much stuff in a batch might wind up going through the same experts. Although we will later on add code to make sure that on average each expert receives roughly the same amount of context vectors, within a given batch things might be imbalanced. To see why that might be a problem, imagine an LLM that is so large that you have different GPUs -- or even different machines -- handling different experts. Something in the trillions of parameters scale, for example. If in your batch, expert i winds up doing most of the context vectors, then the GPU it's on is a bottleneck. The setup I'm describing in this post is sometimes called "token choice" routing. In a sense, each token has chosen -- or, rather, the router has chosen on its behalf -- which experts it "wants" to go to. "Expert choice" routing inverts that; you generate the same logits, but instead of choosing the top- k experts for each token, you choose the top- k tokens for each expert, so that you guarantee balance across the experts. Both of these are something for future investigations, I think :-)  ↩

0 views
Jason Tucker Yesterday

Staying Put with the iPhone 15 Pro Max

Apple held its September event yesterday. New chip, new colors, a redesigned Dynamic Island, and the thing everyone actually cared about: Apple's first foldable phone, officially called the iPhone Duo. I watched the whole thing with my iPhone 15 Pro Max sitting on the desk next to me, and by the end I'd already made up my mind. I'm sitting this one out. This isn't a review. I haven't touched any of this hardware and won't pretend otherwise. It's more of a note on why the upgrade button stayed unclicked this year, and why that's a different decision for me now than it used to be. For a long stretch, a new iPhone showed up in my house every September without much debate. Better camera, better chip, and I'd justify it by pointing at whatever spec sheet Apple put in front of me that year. After a while that was feasible and I switched to the every other year plan with T-mobile and got tired of that too. That stopped once my kids got older. When they were small, the camera mattered. Every version got meaningfully better at low light, at video, at capturing a toddler who wouldn't sit still for a photo. That was worth chasing. Now my wife and I are empty nesters. There's no little kid to chase around the yard, no travel team game where I'm elbowing for a spot on the sideline with a phone up. The upgrade math changed without me deciding anything. It just stopped adding up. Add in some grand kids and I'm back to having the best camera all over again. We're not there yet in this stage of life. My current phone wasn't a random stopping point either. The iPhone 15 Pro Max was the first iPhone with enough RAM to run on-device Apple Intelligence when Apple turned it on. Everything before it physically couldn't do the job, no matter how many software updates rolled through. That felt like a real line in the sand, not a marketing one, and it's a big part of why I stopped here instead of two generations earlier. It's held up. iOS 27 still lists the 15 Pro Max as fully supported, on-device AI features included. The chip hasn't slowed down in any way I can feel, and the camera still does everything I ask of it. That said, today's keynote put the first real crack in that streak. The new Audio Intelligence features on Apple Watch Series 12 and Ultra 4, Live Rewind and Siri Recap, need an iPhone 16 or later to work. Live Rewind pulls back the last 15 seconds of speech to a text snippet on your wrist, and Siri Recap runs in the background all day, summarizing conversations so you can jog your memory later. Neither one runs on a 15 Pro Max, watch or no watch. Not a dealbreaker on its own, but it's a reminder that "fully supported" and "gets every new feature" aren't the same promise, and that gap is only going to widen from here. As an aside, this Live Rewind sounds like a HIPAA nightmare, enjoy. For what it's worth, the watch side of that equation is moot for me anyway. My Apple Watch is a Series 5, several generations older than what these features require, it runs out of battery pretty fast, it does what it needs to do and I'm ok with that for now. Chasing Live Rewind and Siri Recap would mean replacing two devices, not one, which makes the whole feature easy to shrug off for now. None of that means I'm immune to a good pitch, so I did look at what shipped today. The iPhone Duo starts at $1,999 for 256GB, before AppleCare, before the case you'll want because a folding phone with a hinge is not something you drop on tile without consequence. Apple's foldable Pencil ships separately later this year, which says plenty about who this device is actually for. That $1,999 number is also a little misleading on its own. 256GB isn't enough to start a $2,000 phone at, not for something built around two 48-megapixel cameras and video that eats storage fast. My own 15 Pro Max is 512GB, and that's the real floor for me, not the marketing tier. Step up to the Duo's 512GB option and you're at $2,199. A 1TB model, which is honestly where a phone like this should live, runs $2,599. The headline price is the one nobody serious about using the cameras will actually pay. Its so expensive that the tax apps are now sharing how you can write off the phone, or at least portions of it. I get why the engineering is impressive. I don't get paying a premium for a device that, by every early report, still compromises on Face ID in favor of Touch ID because there's no room left for the sensor array, and is arguably worse at the one job my current phone already does well. The camera setup is the part that actually rules it out for me. The Duo ships with two rear cameras and no telephoto lens. I use the telephoto on my 15 Pro Max constantly, enough that losing it isn't a minor tradeoff I'd adjust to. It's a step backward dressed up as Apple's most advanced iPhone. Then there's the fact that this is a first-generation product doing something Apple has never shipped before. A hinge that has to survive years of folding, a crease that has to stay invisible enough to not bother anyone, software that has to handle two very different screen shapes without feeling bolted together. Apple usually gets a device close to right by the second or third try, not the first. I'd rather let other people find the failure points on a $2,000 phone than volunteer mine. The battery is the one real issue on this phone. It's down to 75% capacity with over 1,000 charge cycles, against roughly 600 on my wife's iPhone. Same phone age, mine has clearly done more work on the charging front. A hundred dollars gets it replaced through Apple, which buys time without buying a new phone. But fixing the battery is step one, not the whole plan. The real plan is to save through this next year and replace both my phone and my wife's at the same time, once we've had a full cycle to watch what happens with component prices. Right now the entire industry is dealing with what's being called a memory shortage, driven mostly by AI data centers buying up DRAM capacity that used to go to phones, laptops, and everything else with a chip in it. Analysts are calling for smartphone prices to keep climbing through 2026, and some estimates have the shortage running into 2027 or 2028 before production catches up. That's likely a real part of why the Duo and the 18 Pro line both came in higher than past generations. So the question I'm actually sitting with isn't just camera specs or hinge durability. It's whether this is a temporary spike that eases once memory production catches up, or whether $1,999 phones and $1,199 base models are just the new normal going forward. If prices come back down once the shortage clears, waiting a year saves real money, maybe. If this is the new baseline, waiting doesn't cost me anything either, since the battery swap buys time regardless. There's also a version of this where the market doesn't wait for prices to drop at all. It just leans harder into leasing. If people won't pay $1,999 up front but will pay $58 a month and never think about the total, that becomes the answer the industry settles on instead of cheaper hardware. Watching which way that goes over the next year is as much a part of the plan as watching my own battery health. A while back I also switched carriers, from T-Mobile to US Mobile. That move alone took our service cost from a monthly bill down to a single annual payment that covers the whole year, and it comes out to a fraction of what T-Mobile used to charge every month for the same lines. Once you've broken that habit on the service side, leasing the hardware on top of it stops making sense. Apple offers exactly that now through Apple Upgrade, its Klarna-backed leasing program, and US Mobile has its own version of the same deal. Pay monthly, hand the phone back or buy it out later, repeat every year or two. That's the same recurring payment treadmill I just got off of on the plan side. Trading a small annual number for wireless back into a monthly number for the phone itself would be walking backward, not upgrading. Owning the phone outright and running it five to seven years pairs naturally with paying for service once a year instead of monthly. Both decisions point the same direction: fewer bills showing up, less money leaving on autopilot, and a lot less reason to care what Apple announces every September. Apple doesn't publish a hard number, but the track record is public. Apple's own compliance filing for the iPhone 15 lineup, required under UK law, commits to a minimum of five years of software updates from first sale. The real-world pattern runs longer than that minimum. The iPhone 11 launched in 2019 and is still on the iOS 27 compatibility list this year, which puts it at seven years and counting. Security-only patches have gone even further back, reaching iPhones from 2015 in some cases. Going by that pattern, my 15 Pro Max, released in 2023, still has years of update life left in it even on the conservative end. That's exactly why the plan makes sense: there's no support cliff forcing a decision this year or even next. The battery swap covers the gap, and by the time my wife and I actually pull the trigger on new phones, we'll know a lot more about whether memory prices settled down or just found a new floor. That's the way of life now. Not chasing a keynote every September, but picking the year that makes sense and letting the phone, the market, and the update list tell us when that year has actually arrived. Are you upgrading or waiting another year? Let me know in the comments. Can't comment? Signup for a free account here and you can comment on all of my posts. It keeps the spammer out and makes this place more enjoyable for everyone.

0 views
Kev Quirk Yesterday

2026-09-10 13:49: The constant outrage on #Mastodon (and maybe other social sites?) for is so tiring.

The constant outrage on #Mastodon (and maybe other social sites?) for is so tiring. Thanks for reading this post via RSS. RSS is ace, and so are you. ❤️ You can reply to this post by email , or leave a comment .

0 views
Stratechery Yesterday

The iPhone Duo, The Intelligent Personal Hub, Apple Watch Audio Intelligence

Apple once again demonstrated the power of integrating hardware and software, but it's biggest AI blindspot might be its belief in the primacy of apps.

0 views
Unsung Yesterday

“A great look at the destructive power of an uncaught Infinity value”

There are two distinct flavors of speedrunning: Watching the first one feels like watching the Olympics. The second one is absolutely incomprehensible without someone narrating over and explaining things. This video by Marblr is such narration and explanation of a particular glitch in Portal that significantly shortened a few speedrunning moments: = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/a-great-look-at-the-destructive-power-of-an-uncaught-infinity-value/yt1-play.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/a-great-look-at-the-destructive-power-of-an-uncaught-infinity-value/yt1-play.1600w.avif" type="image/avif"> It’s an hour-long video, and I think it’s a really well-made one, and one you can learn a lot from. The video has some nice history of strafing – similar in some way to Scroll Lock – and motor memory accommodations. It also shows the source code of Portal to explain things. This is fun on its own, but on top of that, Marblr layers some really cool visualizations – watching how one value “infected” by infinity proceeds to infects the others is riveting. Even the levels and their geometry are presented in a really beautiful and informative way, too: I’m still not fully onboard with this branch of speedrunning – if the entry point to your effort is changing mouse sensitivity to an abnormally high value, or if you have to pick a very specific version of the game that just happens not to have patched the glitch you care about, then perhaps things have gotten too esoteric and maybe even academic? But at the very least we get videos like this one – or one I linked to in May – going under the hood of software in some particularly enthralling ways. completing a game as fast as possible through feats of extraordinary mastery, timing, and precision, completing a game as fast as possible by finding all sorts of bugs and exploits in the game (commonly known as “glitches”) that allow the player to find eerie shortcuts between areas without having to play them.

0 views
Justin Duke 2 days ago

It's just a model

I, like many people, have witnessed and felt the whiplash between the seemingly-orgastic highs of the Astra launch and the experience of using the model. Whether intentional or otherwise, a lot of airtime was given in the brief pre-release hypecycle to Astra's ability to generate things that looked cool . Lots of games, renders, and examples cluttered my various timelines. This makes sense: 3D stuff is cool; most people don't really have a sense of the toolchain involved in things like this, and how much of the effort is on the part of the LLM versus the structures it builds upon is somewhat obfuscated; platforms like xtheeverythingapp.com tend to reward videos and animations. I try (unsuccessfully) to remain sober about all of this. But the release of the model, coupled with things like the deeper information around the Hugging Face incident and Collusion.wiki, did ignite a bit of a spark in the back of my head — thinking, as always, however briefly, about the chance that things are really different this time. Did we solve software? Then I started using the model. And much like my quasi-review of Fable 5.1 , I found that the model is good! It's hard for me to really compare and contrast it against Fable 5.1 because I didn't do formal comparisons or anything like that; I just tried throwing a couple of problems I was working on against it, and it handled them all well and the things it couldn't handle were generally due to lack of context and agency, not lack of horsepower. It's very smart and capable, but limited by the tools and harnesses and contexts it has access to (words I deploy in the broader sense, not the industry-specific ones.). All of this mirrors almost exactly my sense of wonder followed by disillusion when I started really kicking around Fable for the first time. 1 Though I do maintain that Fable pre-DOD was a slightly different beast, and felt more akin to a real change in how I programmed. Astra is very good at a lot of things; for me, though, it is increasingly extremely competent at tasks that have already been solved by faster, cheaper models. And at least personally, as someone who runs a non-trivial business where the risk of false positives vastly outweighs the attractiveness of sheer velocity, I struggle to find the right balance here between exploring the boundaries of what is possible and remembering that right now my job is no longer that of a technologist but of an industrial designer. I run a small business where we have absolute license to use models at will, so long as we do so judiciously and securely. I talk about them for no economic or personal reason 2 Disclaimer: I am, as of this writing, an investor in Anthropic through an acquisition they made. ; I want to understand (and thereby master) them out of a conviction that they'll improve my business. I bring that up because I found myself vigorously nodding along to this essay from Armin Ronacher , who writes (amongst other things, which I agree with): I'm more and more convinced that all of AI engineering is Neijuan (内卷, meaning curl inwards). In China it describes a system that demands ever more effort and competition without improving output. The way in which it sometimes shows up in the West is the 996 nonsense. The English term for Neijuan is "Involution" from the book Agricultural Involution. Agricultural involution describes the intensification of farming that raises productivity per square meter while leaving productivity per head unchanged. One way of thinking about LLMs (or any technology) is dichotomous: The biggest step function in software engineering over the past few years was not a specific model but Claude Code. Astra and Fable have improved my life less than full support for subagents and worktrees did. I get that it's unfair of me to pretend those two things are discrete! I don't have enough intelligence to understand how raw power acts as a proxy for various other implications outside my industry. But I do — at least for now — have enough experience to understand that right now it's a poor proxy for operational leverage in mine. There is the raw power a given model can generate, and then there's the design and manufacturing of interfaces that can use that model in useful ways.

0 views
Sean Goedecke 2 days ago

They really do think AI might kill everyone

A recent resignation tweet from an Anthropic researcher has everyone talking about the AI apocalypse again. Among other things, he said: The people building AI earnestly believe that it could kill us all by the end of the decade. Many people found it hard to believe that AI researchers think this way. Some explained it as a PR campaign to promote AI regulation, or as self-promotion , or as a way to boost AI company stock prices. Others felt it had to be impossible, because if you really believed this you’d be bombing datacenters instead of posting on Twitter. In fact, not only do many AI researchers 1 seriously believe this, they’ve been thinking and writing about it since the mid-2000s. Eliezer Yudkowsky — the ur-figure for most modern AI safety culture — has been publishing papers since at least 2008 saying that superintelligent AI could destroy all human life. It’s been such a common idea that the AI research community has abbreviated “how likely you think AI is to kill everyone” to “p(doom)” (i.e. the probability 2 of doomsday) since around 2010. I know it sounds very silly if you’re not in the AI bubble. But it really is true, and if you assume there has to be some different motive you’ll be deeply confused by what AI researchers say and do. They truly do believe that there is a reasonable chance that superintelligent AI will kill everyone. This is why AI researchers care so much about “alignment”: building AIs that share genuinely human beliefs and values. If we build a “misaligned” superpowerful AI — an AI with goals that are alien to us — it might sweep humanity away. It could kill everyone deliberately, e.g. to stop us getting in the way of some goal. It could kill everyone in passing, e.g. like we might pave over an anthill to build a road. Either way, everyone dies. Okay, but how? What do these people think is actually going to happen? AIs are computer programs running in a datacenter somewhere. How could they possibly cause the extinction of humanity? Wouldn’t someone just turn them off? Unsurprisingly, AI research nerds have come up with some concrete answers to this in the last two decades. Here they are, in order of plausibility: An AI could make and release some super-pathogen or virus. In the hope that current AIs can achieve huge breakthroughs in medicine similar to the ones they’ve already achieved in mathematics, we’re setting up autonomous labs 3 . Bioweapons have been terrifying biologists for decades, and with good reason. Historical pandemics have killed up to 80% of affected human populations, and tend to be defeated by accident: the disease happens to evolve into a less virulent strain, or short incubation periods mean that infected people can’t carry the disease far, or some percentage of the population is naturally immune. A plague designed to be maximally fatal — or several plagues in quick succession with different characteristics — could be much worse. Alternatively, an AI could trigger global thermonuclear war . We’re already seeing AI be integrated into military and government decision-making processes . If a rogue AI managed to set off a bunch of nukes, or to coordinate 4 with other countries’ rogue AIs to nuke each other, we could be in an ordinary nuclear apocalypse scenario: billions dead in the strikes, billions dead in the ensuing famine, and so on. It’s commonly assumed that a post-global-thermonuclear-war Earth would still support some tiny human population, but a determined AI could surely find some way to mop up the stragglers. There are some other theories. Once robotics has permeated the world economy, an AI could take over the robots (including drones) to kill everyone, like in Terminator . Or AIs could take advantage of nanotechnology to create self-replicating machines that turn the world into “grey goo” 5 . Or AIs could terraform the planet so as to make it unliveable for humans (as in Nick Bostrom’s famous paperclip example ). Or they could do something else that our puny human brains aren’t able to think of. One common counter-argument 6 here is to say “well, it’d be impossible to extinguish all human life — what about undiscovered tribes in the Amazon, or survivors living in the ruins of modern-day cities?” I don’t know, man. At some point you’re just conceding the argument: the policy positions you’d adopt if you thought AI might wipe out 99% of humans are the same as if you thought it might wipe out 100%. And like I said above, if an AI can kill almost everyone, it’s probably smart and capable enough to finish the job somehow. Another is to say that the government will simply step in and nationalize the AI labs when the situation gets too dangerous. Maybe! But this kind of concedes the argument: a technology important enough to be fully taken over by the government is a terrifyingly dangerous technology. A third is to say “well, someone would just turn it off”. I don’t find this plausible at all: an AI powerful enough to build a super-plague is an AI sophisticated enough to pretend it’s curing cancer, or to exfiltrate itself to some datacenter where it won’t be turned off, or to take some other countermeasures. Why would you work in AI, if you believe this? Why wouldn’t you go live in the woods somewhere, or start bombing datacenters , or assassinating AI lab CEOs? For a few reasons. An AI powerful enough to end humanity is an AI powerful enough to save it. I wrote about this in Help peer : many AI researchers believe that the only way for humanity to truly survive long-term is with the help of superintelligent AIs, so long as someone can figure out alignment. Isn’t this a huge risk? Maybe not. If somebody is going to build superintelligent AI, you might be obligated to try and do it first. You can’t go and bomb every datacenter in the world, after all. Why does it matter who’s first? Some popular theories of AI development involve a “foom” or “hard takeoff”: the first time someone really cracks self-improving AI, capabilities will increase exponentially, because smart AI will be better able to make itself smarter, ad infinitum 7 . There are no draws in the AI race . The first lab to figure out smart human-like intelligence will be the first one to figure out wildly superhuman intelligence, and thus will be in a position to stop anyone else from doing it. This is an under-discussed point in the AI risk debate. Lots of AI researchers believe that the first thing a true superintelligence will do is reach out and stop all other AI research : either by hacking the labs, persuading them to stop, or literally drone-striking their datacenters . According to this view, if you’re an AI researcher and you think you can build an aligned AI, you should be working 24/7 so you can manifest God, and you should wake up every morning gripped by the fear that someone elsewhere has manifested the Devil, and your training datacenter no longer exists. I have been on the fringes of this world for my entire adult life. I read Meditations on Moloch as a young adult and wanted to get into AI. I am one of the few people to read the entirety of the Sequences , Eliezer Yudkowsky’s million-plus-word magnum opus about rationality. I was too young for the Extropians mailing list, but I’ve spent years on LessWrong . On the other hand, I’m not a card-carrying rationalist: I think if you have a strong intuition on one side and a convincing-sounding argument on the other, you should pick the intuition 8 . I’m a deontologist , not a utilitarian. I don’t even live in San Francisco! I’m conflicted about AI risk. The current behavior of AI agents does seem to vindicate a lot of the early science-fiction-sounding worries of the AI doomers, but modern LLMs are a lot more human-like than the alien minds in the apocalypse scenarios, and in general it does just seem too silly to credit (I guess I’m picking the intuition here). However, it bothers me to see people dismissing these people as part of a PR operation, or as liars looking to boost an upcoming AI lab IPO, or as isolated crazies who haven’t thought their position through. Whatever else you say about the AI doomers, they have more than two decades’ history of explicitly spelling out exactly what they believe and why, even when it was complete science fiction to talk about AI at all. They’ve earned the right to be treated as sincere. In this post I’m going to use “AI researcher”, “AI safetyist”, and “rationalist” as reasonably synonymous terms for “someone who thinks there’s a chance AI kills everyone”. You could write a whole other post about the relationship between the rationalist/“AI safety” community and making concrete numerical predictions for unlikely future events. Alternatively, smart enough AIs might trick scientists with ordinary labs to produce dangerous substances. This is beyond the scope of this post, but many AI researchers believe that super-smart AIs will inherently come to agree with each other and eventually to coordinate without ever having to communicate, simply because they can predict what the other one will do. This was the most popular theory in the late 2000s, when nanotechnology was trendier. Relegating this counter-argument to a footnote because I hate it: many people say “we shouldn’t worry about AI risk, because climate change (or AI misinformation, or some other thing) is more urgent and serious”. You simply do not have to choose: it is possible to worry about multiple risks at the same time. Some people advocate for a “slow takeoff”. However, the main proponent of that view is Paul Christiano, who has just today joined OpenAI, citing “a meaningful risk that rapid acceleration in AI capabilities leads to catastrophic and irreversible loss of control in the very near term”. This is kind of a Michael Huemer-ish position in epistemology. In this post I’m going to use “AI researcher”, “AI safetyist”, and “rationalist” as reasonably synonymous terms for “someone who thinks there’s a chance AI kills everyone”. ↩ You could write a whole other post about the relationship between the rationalist/“AI safety” community and making concrete numerical predictions for unlikely future events. ↩ Alternatively, smart enough AIs might trick scientists with ordinary labs to produce dangerous substances. ↩ This is beyond the scope of this post, but many AI researchers believe that super-smart AIs will inherently come to agree with each other and eventually to coordinate without ever having to communicate, simply because they can predict what the other one will do. ↩ This was the most popular theory in the late 2000s, when nanotechnology was trendier. ↩ Relegating this counter-argument to a footnote because I hate it: many people say “we shouldn’t worry about AI risk, because climate change (or AI misinformation, or some other thing) is more urgent and serious”. You simply do not have to choose: it is possible to worry about multiple risks at the same time. ↩ Some people advocate for a “slow takeoff”. However, the main proponent of that view is Paul Christiano, who has just today joined OpenAI, citing “a meaningful risk that rapid acceleration in AI capabilities leads to catastrophic and irreversible loss of control in the very near term”. ↩ This is kind of a Michael Huemer-ish position in epistemology. ↩

0 views

Intelligence and Computation

This article has been posted to my Substack also. There are two kinds of mathematics: “simple” ones, with which you can create a complex video game, implement a working replacement for a worldwide financial ecosystem (Bitcoin), or even raw intelligence, with neural networks. This math is simple in the sense that it is basic linear algebra (for video games and AI) and even simpler (but extremely convoluted) math, for Bitcoin. What makes this type of math powerful is that when you combine it with a computing substrate, you get very impressive and useful results. You can simulate reality, get conversational intelligence, etc.

0 views
Farid Zakaria 2 days ago

Review a pull request by booting it

tl;dr; trynix-preview is a GitHub action that comments a link on a pull request which lets you boot the PR’s build in the browser using https://trynix.dev . No servers, just browsers. I ended my earlier trynix post with a list of ideas I think we could accomplish now that we can boot arbitrary paths in the browser. The most obvious one was to let a reviewer boot a pull request’s build in the browser for testing, validation and feedback. That is now real. 🤯 A demo is worth a 1000 words: here is a pull request ( PR#31 ) against my sqlelf project, from a fork , with the comment our action left on it: Click the link and a Linux machine boots in your tab with that PR’s on . You did not clone anything, you did not build anything. No servers, no SSH, no VPN, no Docker, no VM, no cloud. Just a browser and a link. 😈 As with any GitHub action, it’s just a few lines to add to your workflow. The caveat is that you must have built and cached the path already, so the action can link to it. The action does not build or cache anything. The action publishes and builds nothing. Whatever already fills your cache keeps doing it, and the action’s whole job is to simply provide the store paths via and hand the cache’s URL and public key to the browser. It is not Nix cache provider specific, but I do recommend Cachix because it is free for open source up to 5GiB. 1 You can checkout my trynix.yaml workflow for the full example. You have to set in the step because the workflow runs on a fork’s pull request, and that has security implications. 2 If that is not your cup of tea, there is a version where a maintainer types on the pull request which kicks off the workflow. In either case, the workflow runs on the default branch and checks out the pull request’s code, so a fork cannot edit the workflow that builds it. Did I just upend all CI products by easily letting reviewers boot a PR? Unfortunately, no. 🥲 The performance for large binaries is pretty bad. Even with many of the improvements I AI-assisted into the engine, large binaries can still take 1-2 minutes to execute. 3 Nevertheless, this is still a pretty amazing workflow and showcases the power of Nix. Maybe as we get closer to AGI, our AI overlords will be able to optimize the engine to execute large binaries in a few seconds, but for now, the action is best suited for small to medium-sized binaries. You should definitely sign up for Cachix but you can test this out without it since the free tier is very generous.  ↩ I recommend a private segregated cache for pull request builds, so that a fork cannot push to your main cache.  ↩ I added a benchmark page, https://trynix.dev/bench/ , to the site with a lot of rich data on boot and run times for various applications.  ↩ You should definitely sign up for Cachix but you can test this out without it since the free tier is very generous.  ↩ I recommend a private segregated cache for pull request builds, so that a fork cannot push to your main cache.  ↩ I added a benchmark page, https://trynix.dev/bench/ , to the site with a lot of rich data on boot and run times for various applications.  ↩

0 views
iDiallo 2 days ago

We own the Glass

This was so casually uttered by LG executives in their sales pitch to advertising companies. We own the glass. What they mean is, they own your tv. They own everything on it. You may have paid for the box, but they own the glass. They can run anything they want on it, and you have agreed, in the terms and conditions, that you will inform anyone entering your household that the TV may record them without warning. TVs come equipped with microphones, wifi radio, and tracking software. That’s enough to monitor everything that’s happening around it. If you haven't watched the research video essay published by Gamer Nexus , I highly recommend it. I've said it many times before, TVs are cheap because you aren't simply watching them, they are watching you and listening, and reporting back to the mothership. This info is then used to sell you more stuff. What scares me in all this is that the solution I often tell people may no longer be true. I often tell people, just disconnect your TV from the network and you are all set. Now imagine just for a second. What if they add a cellular radio to the device? One with an esim. Then that is it. You can no longer disconnect your device from the internet. They now own the glass in your house, and have a permanent connection that you cannot disable. They own the glass. We are merely vessels to their business.

0 views