Posts in Data (20 found)
ava's blog Yesterday

link dump - catching up on my online reading

While I am slowly getting back up on my feet after a tough time, I am catching up on emails (will still take a bit!), my RSS feed reader, and several newsletters that have accumulated in my inbox. Here's what I picked out to read and share: This Is Capitalism: Apple's Hidden Data Workers at the Shadows of the AI Boom - 19 page paper detailing what data workers actually do and what problems they deal with, written anonymously by a data worker interviewing their colleagues. Parts of the work descriptions remind me of the work in Severance . Is this really a good reason to triple datacentre capacity in Europe? - online blog post tracing where the idea to triple the EU's data centers comes from that is mentioned in several EU AI strategies. Turns out it's from be a blog post from Savills, a commercial real estate firm who profit off of data centers being built. Fake US thinktank set up and funded by Israel sought to game AI for propaganda - AI slop meant to absolutely flood the web to be included in AI training funded by the Israel government to spread disinformation in AI answers. A new force is increasing inequality in America - WaPo article about how AI is not leveling the playing field or bridging the gap between poor and rich, but instead worsening the gap. People making the most use of AI are concentrated in richer urban areas and are already often rather wealthy, while the AI data centers are in poorer neighborhoods and the data workers are often migrants or in the Global South. Rich people can invest into AI and its stock, therefore profiting off of the hype and concentrating even more wealth. An operational framework for AI literacy in the workplace - 15 page paper from Interface EU addressing the vagueness of "sufficient AI literacy" that is often mentioned in EU AI legislation. It proposes a cumulative three-tier framework based on the nature and consequences of a worker's interaction with AI which then dictates the level of literacy required and how to attain/ensure it and measure it. The Quiet Erosion of Collective Action Under Digital Surveillance - article on chilling effects of permanent surveillance and the feeling of constant suspicion which continues to erode activism. Flipping the kill switch: I survived 72 hours without US tech - online article about an experiment to live without reliance on US tech. The US has such a strong monopoly that almost all online services and tech are unusable with this rule. Gone in one click - assessing the socio-economic impact of browser-level consent in Europe - small informative flyer style PDF by the Implement Group showing figures about cookie consent rates and ad industry revenues depending on the mode of consent. Inside the growing vigilante movement to knock out Flock surveillance cameras - online article by The Guardian. I admire these people, and we need more civil disobedience now, everywhere. Not just against Flock; against Meta glasses wearers, against Ring camera owners (Yes, you too! None of you are the "good ones" with "valid" reasons!) and more. There's many ways to affect these devices that you can find online or just get creative with it. Keep yourself safe, don't write about it, don't record yourself doing it, don't discuss it via digital means, leave your devices at home, and leave no fingerprints. Did someone wearing Meta Glasses film you today? Are you sure? - another Guardian online article, this time about the spy glasses and the people who enable hiding the recording light on them. The man behind GhostMeta is actually so vile and disgusting; anything else I could say would violate the Code of Conduct this blog is hosted on. German links: HeißeLuft.org - German website with interactive map showing where AI data centers are planned, in development, and paused, together with information on protests. Made me discover that they are planning on building one not too far away from me... Deutsche Post trainiert ihre KI mit Ausweisfotos - article about how Deutsche Post is training their AI with ID pictures they get via digital identification procedures. It's not voluntary as they claim, as you need to give permission before being allowed to proceed. Verhaltensscanner in Berlin: Harte Kritik an der KI-Überwachung - online article about the new camera installed in Berlin that will analyze all people in that area via AI surveillance software by Adesso, with more cameras to follow. They want to put them up in high crime rate areas , but keep secret what the standards for this are, which enables a mass roll-out of them if they wanted to without any oversight or control. There has already been one mix-up leading to higher crime reported in an area than actually happened. These cameras already also exist in Hamburg and Mannheim. Möglicher AfD-Sieg in Sachsen-Anhalt - online article detailing the fear of queer people and people of color of an upcoming potential win of the AfD in their state. Afd-Gutachten.de - website containing some stats and a PDF report of a legal assessment on the chances of a successful AfD ban. „Gipfel gegen Linksextremismus“: Mit Trump gegen die Antifa - online article about the cooperation of Germany with the US on its fight against antifascism. Unfortunately it has continued, with the German government realigning to focus more on supposed "leftist extremism" and even re-distributing money away from leftist projects, which mostly hits projects aimed at helping queer people and migrants. Wie weit ist Deutschland beim digitalen Gewaltschutz? - an online article about the really embarrassingly low standards of protection against digital violence, especially image-based ones like deepfake nudes and revenge porn, in Germany. Lots needs to be done in general, but especially to even meet the new EU standards. Wer ist für Straftaten der KI verantwortlich? - legal article about the criminal liability of autonomous AI in Germany, and how crimes done by AI agents are pushing the legal system to its limit as we only legislate for humans. Published 29 Aug, 2026 This Is Capitalism: Apple's Hidden Data Workers at the Shadows of the AI Boom - 19 page paper detailing what data workers actually do and what problems they deal with, written anonymously by a data worker interviewing their colleagues. Parts of the work descriptions remind me of the work in Severance . Is this really a good reason to triple datacentre capacity in Europe? - online blog post tracing where the idea to triple the EU's data centers comes from that is mentioned in several EU AI strategies. Turns out it's from be a blog post from Savills, a commercial real estate firm who profit off of data centers being built. Fake US thinktank set up and funded by Israel sought to game AI for propaganda - AI slop meant to absolutely flood the web to be included in AI training funded by the Israel government to spread disinformation in AI answers. A new force is increasing inequality in America - WaPo article about how AI is not leveling the playing field or bridging the gap between poor and rich, but instead worsening the gap. People making the most use of AI are concentrated in richer urban areas and are already often rather wealthy, while the AI data centers are in poorer neighborhoods and the data workers are often migrants or in the Global South. Rich people can invest into AI and its stock, therefore profiting off of the hype and concentrating even more wealth. An operational framework for AI literacy in the workplace - 15 page paper from Interface EU addressing the vagueness of "sufficient AI literacy" that is often mentioned in EU AI legislation. It proposes a cumulative three-tier framework based on the nature and consequences of a worker's interaction with AI which then dictates the level of literacy required and how to attain/ensure it and measure it. Great quote from it: "No evidence yet shows that these trainings work, and three gaps might explain the reason. The first is motive. Corporate training aims at productivity and teaches people to use the tools well, whereas the law cares whether operators understand how systems fail and cause harm. A workforce fluent in prompting can still be illiterate in the sense a regulator means: trained to produce good output, but not to recognise when a model misleads or to know its duties under data-protection and risk rules." The Quiet Erosion of Collective Action Under Digital Surveillance - article on chilling effects of permanent surveillance and the feeling of constant suspicion which continues to erode activism. Flipping the kill switch: I survived 72 hours without US tech - online article about an experiment to live without reliance on US tech. The US has such a strong monopoly that almost all online services and tech are unusable with this rule. Gone in one click - assessing the socio-economic impact of browser-level consent in Europe - small informative flyer style PDF by the Implement Group showing figures about cookie consent rates and ad industry revenues depending on the mode of consent. Inside the growing vigilante movement to knock out Flock surveillance cameras - online article by The Guardian. I admire these people, and we need more civil disobedience now, everywhere. Not just against Flock; against Meta glasses wearers, against Ring camera owners (Yes, you too! None of you are the "good ones" with "valid" reasons!) and more. There's many ways to affect these devices that you can find online or just get creative with it. Keep yourself safe, don't write about it, don't record yourself doing it, don't discuss it via digital means, leave your devices at home, and leave no fingerprints. Did someone wearing Meta Glasses film you today? Are you sure? - another Guardian online article, this time about the spy glasses and the people who enable hiding the recording light on them. The man behind GhostMeta is actually so vile and disgusting; anything else I could say would violate the Code of Conduct this blog is hosted on. HeißeLuft.org - German website with interactive map showing where AI data centers are planned, in development, and paused, together with information on protests. Made me discover that they are planning on building one not too far away from me... Deutsche Post trainiert ihre KI mit Ausweisfotos - article about how Deutsche Post is training their AI with ID pictures they get via digital identification procedures. It's not voluntary as they claim, as you need to give permission before being allowed to proceed. Verhaltensscanner in Berlin: Harte Kritik an der KI-Überwachung - online article about the new camera installed in Berlin that will analyze all people in that area via AI surveillance software by Adesso, with more cameras to follow. They want to put them up in high crime rate areas , but keep secret what the standards for this are, which enables a mass roll-out of them if they wanted to without any oversight or control. There has already been one mix-up leading to higher crime reported in an area than actually happened. These cameras already also exist in Hamburg and Mannheim. Möglicher AfD-Sieg in Sachsen-Anhalt - online article detailing the fear of queer people and people of color of an upcoming potential win of the AfD in their state. Afd-Gutachten.de - website containing some stats and a PDF report of a legal assessment on the chances of a successful AfD ban. „Gipfel gegen Linksextremismus“: Mit Trump gegen die Antifa - online article about the cooperation of Germany with the US on its fight against antifascism. Unfortunately it has continued, with the German government realigning to focus more on supposed "leftist extremism" and even re-distributing money away from leftist projects, which mostly hits projects aimed at helping queer people and migrants. Wie weit ist Deutschland beim digitalen Gewaltschutz? - an online article about the really embarrassingly low standards of protection against digital violence, especially image-based ones like deepfake nudes and revenge porn, in Germany. Lots needs to be done in general, but especially to even meet the new EU standards. Wer ist für Straftaten der KI verantwortlich? - legal article about the criminal liability of autonomous AI in Germany, and how crimes done by AI agents are pushing the legal system to its limit as we only legislate for humans.

0 views
matduggan.com 2 days ago

You Know GDPR Is Good Based on Who Hates It

FDR has always been one of my favorite presidents, second maybe to Lincoln. Both were men the establishment assumed were one of them until, to their horror, they governed like they weren't. Both could trash the opposition in one breath and take the moral high ground in the next. One of my favorite Roosevelt lines growing up came from Madison Square Garden, October 1936, standing before a crowd that included plenty of people who wanted him dead: What I love is that he doesn't argue with the hate. He doesn't say they're wrong to hate him, or that the hate is unfair. He says the hate is evidence. The process is working. I've always treated it as a metric: if you're doing something hard and nobody hates it, you probably aren't doing it. If the right people hate it and those people happen to be some of the worst people alive, so much the better. By that standard, the GDPR (Europe's General Data Protection Regulation) is doing beautifully. It is impossible to go anywhere in a technology space online without hitting a wave of commentary about how stupid GDPR is. It was written by bureaucrats who don't understand the amazing potential of unrestricted technology. These US-based critiques almost always lean on the oldest trick in cyberlibertarianism: we don't have time to regulate, we must simply adapt and ride the wave. Nobody has time for government. Of all GDPR's consequences, none gets more attention than the cookie banner, which critics present as the inevitable result of government meddling. Blaming GDPR for the cookie banner is like blaming the health inspector for the roaches. The banner is deliberate vandalism, a dark pattern engineered to exhaust you before you can learn anything about the surveillance apparatus humming behind the "OK." Ironically the banner designed to hide the machine has taught the public more about the machine than a thousand podcasts ever will. Even non-technical people stop at "your data is shared with 996 partners." "For a sports score website?" So why is the tech commentary community so loud about this? Because they understand what's at stake. If consent must be freely given and easy to refuse, the industry's power shrinks exponentially. People might decide who has their data, how long it's kept, and what it was collected for. You can only imagine how that thought keeps a Meta executive up at night when he's not eating endangered animals, or ignoring calls from his children whose names he has forgotten while on a tacky yacht. Consider the following example. Did you know Google was doing this every single time you searched on Google ? Did your dad? So even in the most maliciously compliant form the regulation does provide value and information. Remember the hatred is the metric. How did we get here, where US tech companies end up regulated by Brussels? Why isn't the US government regulating US corporations anymore? If the rules are so terrible, why did nobody choose market exit? Has the EU become the world's "privacy cop" or, in the inverse, the biggest player to protect a fundamental human right to privacy? Ah who doesn't remember where they were on GDPR Day. Since we're all socialists in the EU, we stood up from our government issued desks and gave the required three cheers for regulation, then resumed being on vacation for 6 weeks. Obviously after stopping by my free doctor on my way to the airport. On May 25th, 2018, GDPR took effect to the sound of American commentary, the way fireworks take effect to the sound of dogs. You can tell American CEOs were aware of the regulation based on the speed by which they copied the language from it. Right before it took effect Brad Smith, the president of Microsoft, tweeted "We believe privacy is a human right." Tim Cook was right behind, telling CNN that "privacy is a fundamental human right". The framing of privacy as a human right is one of the key elements of the EU approach with GDPR. This is in stark contract with the US legal system which views information privacy as more of a market problem. You are all informed individuals in the wide marketplace of data exchanges and are left mostly to your own devices. In theory there should be regulations by the US of things like unfairness, deceptions and other market failures but in practice that doesn't happen. Europe is no stranger to this fight. The German state of Hesse passed the world's first data protection law in 1970, also known as the year the Beatles broke up, and set a standard we still fail to meet today: At the launch of GDPR there were 126 countries with data privacy laws of some sort. What you see with this sea of legislation is an overwhelming consensus that what GDPR was attempting to do was correct. In fact you see a pretty high level of global convergence of standards. All 126 laws descend from the same commandments the OECD carved in 1980: collect only what you need, say what it's for, keep it safe, let people see and correct it, and don't be a creep about any of this. Fifty years later, the American internet industry is still stuck on commandment one. So first the often-repeated sentiment that this is a flight of EU fancy is straight up incorrect. Something you could describe as the "European standard" for data privacy quickly became a global standard. Anu Bradford calls it the Brussels Effect: Europe regulates, the world complies, because despite American bluster, Europe is a market nobody can leave. In the US financial market, companies who surrendered the EU market would have quickly found themselves with new CEOs as their previous leaders suddenly discovered health problems or a deep love of their families that they had ignored for years. Especially given the increasingly frosty relationship between the US and China, companies that were pushed out of China due to regulation and increased domestic competition cannot lose the EU market. It's also difficult to screen a lot of services for EU customers, especially because the laws follow the personal data of EU residents whenever and wherever the information is transferred outside of the EU. Think of it like trying to sort luggage based on what stickers are on the outside. The EU is also set up in such a way where enforcement of such a law becomes possible. Every member state has a Data Protection Authority, who are charged with assisting individuals in protecting their rights, advising domestic legislatures on the functioning of existing regulation and finally enforcing the law. The EU is also different from the US in that it is open to exploring precautionary regulatory action. We see this with the EU Artifical Intelligence Act which tried almost immediately to get some controls on the industry right at the beginning. Link So the combination of a robust option for enforcement combined with an increased appetite for regulation in general and a high level of respect for personal privacy made the EU the logical source for this legislation. Want to know what the teeth look like? In 2011, an Austrian law student named Max Schrems asked Facebook for everything it had on him and got back 1,200 pages, much of it stuff he'd never volunteered. He filed a complaint from a dorm room. Four years later, the Court of Justice of the EU had voided Safe Harbor, the transatlantic data treaty, on the strength of it. A college student complaint killed an international agreement signed by presidents and prime ministers....multiple times. The best proof of GDPR's power is what it did to Japan. On January 23rd, 2019 the EU and Japan reached a deal which allowed for the free flow of personal data between the two economies. It established an overarching privacy law with a core set of individual rights and enforcement by independent supervisory authorities. The process took 2 years, which isn't a surprise because before this process Japan had very weak, swiss-cheese regulations. In 2014 Graham Greenleaf chose the title "The Illusion of Protection" for his chapter about Japan in an overview of Asian privacy laws. The private sector was basically unregulated, it has "easily manipulated exceptions" to its rules concerning the use and disclose of personal data, its absence of provisions for sentivie information and had no restrictions on data exports. It was nowhere near the level of protections that the EU would expect for information sharing. You can draw a straight line from a bargaining table in Brussels to new rights for a retiree in Osaka. Japanese data brokers hate it, which again is the point . Why hasn't the US sought the same arrangement? Because everyone involved knows the application would be denied. Which raises the real question: why can't the country that invented most of this technology produce a rule for it? Why is the US stuck pretending there's no reason to regulate while the rest of the world moves on? There's no better text on what has happened in the US with the data economy than The Age of Surveillance Capitalism by Shoshana Zuboff. It's a good read and I won't ruin it for you. First, the definition of Surveillance Capitalism from the book. Didn't read that? I don't blame you. In English they found oil and the oil was us. And as us Americans love to do, we immediately discovered a sudden love of freedom in the presence of oil. Effectively here's what happened. The 9/11 terrorist attacks had a ripple effect in US regulations, effectively derailing the momentum that the domestic US regulatory organizations had about starting to build frameworks around personal data. The focus became on security and not privacy. Quickly public intelligence agencies and the fledgling surveillance capitalist business in Silicon Valley found each other and carved out the concept of "surveillance exceptionalism". If you were spying to keep us safe, then it's not spying it's public service. As time went on, the corruption of American politics rendered the possibility of regulation less and less feasible at the federal level. Google and Facebook poured tens of millions into lobbying (for non-US readers, lobbying is a nice way of saying bribery that is legal). There also became an "open door" between government and tech, with 197 people moving back and forth from Washington to the Googleplex. What these companies learned is that by studying our behavioral data during periods of relaxation or play was the most value, allowing them to accurately and reliably push people towards profitable outcomes. US consumers were aware of this to some extent, often repeating phrases like "If it's free, then you are the product". That was true before, but in the new digital economy it is no longer true. We're not even the product anymore, which is why all these companies no longer give a solitary shit about whether their stuff is good or fun to use. We're the raw material that they mine. They know every single thing about us and we don't know anything about them. They accumulate all the data from us but not for us or for our benefit. Behavioral modification though the accumulation and manipulation of this data is now the wealth engine of the United States. Now we are in a market failure scenario. No individual company will disarm first, and any regulator can be captured for what these companies spend on catering. Only law with real enforcement teeth changes the math. And the US cannot pass laws with teeth anymore, not because voters don't want them (polling says they do), but because the pipeline is purchased. When someone says America "can't" regulate tech, they mean it the way a hostage "can't" reach the phone. GDPR didn't spread because Europe is heroic. GDPR spread because Brussels is what fills the room when Washington leaves it. Thanks to the amazing https://decryptads.com you can see what it looks like in real time. Let's take The Verge. This is a relatively straightforward tech commentary website that has a paywall. It should have a very simple supply chain in terms of advertising. What we see is the opposite. There is a giant network of companies trading, selling and bidding on the information from this website. This isn't the fault of The Verge, this is just how the machine operates. Your data is like fish at a fish market. Everyone gets to inspect the merchandise and decide whether they want to buy without you being involved. Ironically we only get this information because of the and which exist to combat advertising fraud. Even a well-run website operated by technologically literate people geared towards tech enthusiasts behind a paywall is not immune to this system. What you see here is that there is no fucking escape . If you want to have a large web presence and have bills to pay, you need to participate. And The Verge is doing it right! 59 declared partners and none of them are actively preparing for war with the US. In the US though this is not a new problem. The economic historian Karl Polanyi came up with the concept of "Double Movement". You can read it here. Think of it like this. Imagine you strip American Capitalism down to a tug of war. In the first round, the businesses go first. This is the "let the invisible hand decide" part. Everything becomes a product you can buy and sell including land, work, even money itself. Think of this like removing all the referees from a game and letting players do whatever they want. Then there is round 2, which is "wait this is hurting people". People demand protections like minimum wage laws, workers rights and trade regulations because the "free market" never seems to regulate itself but is causing real harm. The referees come back, but now with a rulebook written by the players who got hurt. His argument is that a totally free market is a fantasy. It's a fantasy because markets need government support to operate. Even people who pretend they despise government interference rely on them. Copyright and trademark doesn't matter when you need to train an LLM, but it matters a whole lot when I start marketing my iWatch smart watch. That rhythm of abuse, then correction governed American capitalism for a century. For surveillance capitalism, round two never comes. There is no functional federal counterweight as I write this in 2026. Some states are trying, and good for them, but regulating the internet state by state isn't progress it's just stopping the bleeding. Meanwhile the brokers buying and selling your life exist entirely outside your scrutiny, have no fear of comprehensive reform, and can swing a statehouse election for what they spend on a quarter's catered lunches. Entities this powerful cannot coexist with a functional democracy — a fact they seem to have considered, given the rise of the "Nerd Reich". https://www.npr.org/2026/08/10/nx-s1-5925350/the-nerd-reich-tracks-the-unmasking-of-silicon-valleys-true-politics So the order of operations is simple. Washington can't act, so Brussels does. Brussels acts, and the world follows, because the market is too big to leave. That is the whole story of GDPR: not European ambition, but American absence. A sign of the declining empire if you will. Roosevelt gave that speech at Madison Square Garden on the last night of October 1936, and a week later he won forty-six states. The people who hated him got Maine, Vermont, and the next ninety years of being wrong about everything they hated. That's the thing about being hated by the right people which is it is a currency and a valuable one. Privacy regulation will get there too. Not because the industry repents because industries don't, but because this is how every one of these stories ends: seatbelts, smoking sections, lead paint. Normal, then scandalous, then unthinkable. Someday someone will ask what an ad network for children was and refuse to believe the answer. And the executives who fought this will be on a boat somewhere, explaining that they were for it all along. Accept all.

0 views
Martin Fowler 3 days ago

Making Your Data Ready for Agentic AI

Lots of organizations are excited about what AI can do to streamline their processes, save money, and juice margins. But AI's capabilities are founded on the data that AI accesses, and for many organizations that foundation is little more than sand. Pramod Sadalage and Prem Chandrasekaran write about how to build a reliable foundation of data that can be accurate and trusted.

0 views
マリウス 6 days ago

The Cables that Connect the World

At the northeastern edge of La Línea de la Concepción , on a scrubby Mediterranean beach called El Burgo–Torrenueva , there is an old battlement-tower, La Torre Nueva , and not much else. It was part of the system of coastal watchtowers during the 16th century that would defend the area against the incursion of the Barbary corsairs . The coordinates are . Walk the tideline and you would never know that buried two metres beneath the sand, a fibre-optic cable comes out of the sea here and turns into the internet. It’s the start of a line that runs across the Strait of Gibraltar to Ceuta , on the African coast, and on toward two continents. Nearly everything you do online that crosses an ocean passes through a cable like this, ending, in most cases, underneath a similarly unremarkable patch of coast. Note: Ceuta is an interesting place by itself, that has recently gained some attention and that would also make for an interesting write-up of its own. However, the tl;dr is that it is an autonomous Spanish city of some 85,000 people sitting on the North African coast, bordering Morocco , which means the European Union has one of its very few land borders with the African continent running straight through a peninsula most people could probably not even point to on a map. It has been held by the Spanish crown since 1668, it had been Portuguese before that, and Morocco seemingly never stopped claiming it. For our purposes, though, what matters is that the small enclave, until very recently, hung off the mainland’s network by a single ageing link. When we talk about the internet we do so as if it were air. Ambient, ownerless, and everywhere. In reality, however, it is the exact opposite, because international data doesn’t (normally) travel by, let’s say, satellite, despite what most people might assume. It travels through roughly 1.5 million kilometres of very real (and very owned) fibre-optic cable lying on the seabed, surfacing at a small number of carefully chosen landing points. For these landing points you normally need a gently sloping seabed, mild currents, and little marine traffic, so that anchors and trawlers don’t sever the line. Suitable spots are scarce enough that the same beach usually becomes the shared landfall for several cable systems at once. Unlike what you might be thinking of at first, submarine cables aren’t your run-of-the-mill Ethernet or fibre cable. The hardware that does the heavy lifting out in the deep ocean is about as thick as a garden hose with roughly 25mm across and weighing in at around 1.4 tonnes for every kilometre. The part that carries your data is a small bundle of glass fibres, each one around the same thickness as human hair, sitting in the very middle. Everything else wrapped around those fibres is there to keep them alive in a deeply hostile environment. Working outward from the core, the fibres sit in a water-blocking gel inside a thin copper or aluminium tube, which is sheathed in polycarbonate, then an aluminium water barrier, then a layer of stranded steel wires that give the cable its tensile strength, then a wrap of mylar tape, and finally an outer skin of polyethylene. The copper is for power, because the cable doubles as a very long extension lead, which we will get to in a moment. Closer to shore, where trawlers and anchors roam, the whole thing gets one or two further jackets of galvanised steel armour wire, swelling it to 50mm or more in diameter and several times the weight. Hence, the cable that surfaces on our Spanish beach is buried a couple of metres down and not simply left lying on the sand. The reason a copper conductor runs the entire length is that light, no matter how pure the glass, slowly fades as it travels, and so every 50 to 80 kilometres the cable is interrupted by a repeater , which is an optical amplifier that boosts the signal back up before passing it along. Each repeater needs electricity, and because the fish sadly still didn’t manage to install power sockets on the ocean floor, the shore stations at either end have to feed a direct current of anywhere between 3,000 and 15,000 volts down that copper core, to literally power the cable from both ends at once. On top of the amplification, modern systems lean on a stack of clever tricks to keep the signal intelligible across thousands of kilometres of glass, including wavelength-division multiplexing to cram many separate colours of light down a single fibre, coherent detection to read them back out, and forward error correction to repair whatever gets garbled along the way. Length, then, is mostly a question of power and amplification rather than of the glass itself. Shorter hops can dispense with repeaters entirely, hence an unrepeatered span will happily run to around 250 kilometres on amplifiers at each end alone, which is roughly the length of the line we started this post with. At the other extreme, a single system can stretch across an ocean, and the longest of them, like the 2Africa cable encircling the continent it is named after, run to tens of thousands of kilometres. The actual manufacturing and laying of these cables is, perhaps a little surprising for something the entire global economy rests on, the business of only a small handful of companies. The bulk of the world’s submarine cable is built and installed by just four suppliers, namely the American SubCom , the French Alcatel Submarine Networks , the Japanese NEC , and the Chinese HMN Technologies . They own and operate the specialised fleet of cable-laying ships, which aren’t exactly the kind of boat you would recognise from a harbour, but more like a purpose-built vessel carrying thousands of kilometres of cable coiled in enormous tanks below deck, rolling it out over the stern at a steady walking pace as they crawl across the ocean. Deploying a new system is a multi-year effort that begins long before any ship leaves port. First somebody, these days increasingly a content giant rather than a phone company, decides a route is worth having and assembles the money for it, either alone or as a consortium of several owners sharing the bill. Then comes a marine survey, in which a ship maps the intended path along the seabed to find the gentlest, safest route around wrecks, trenches, and other people’s cables, followed by the permitting, which is the paperwork of securing landing rights and concessions from every jurisdiction the cable so much as touches. As we are about to see on the Spanish beach, this can generate a remarkable quantity of bureaucracy . Only once all that is settled does the cable get manufactured to length, loaded onto the ship, and laid, with the vessel simply lowering it onto the seabed in deep water and a sea plough burying it a metre or two beneath the sediment closer to shore, where the danger from fishing and anchors is greatest. A working ship covers somewhere in the region of 100 to 200 kilometres a day, so an ocean crossing takes several weeks at sea. A transatlantic system running some 7,000 kilometres typically costs in the order of 250 million USD, while a longer trans-Pacific route can easily climb towards 400 million, and the cable itself runs anywhere from roughly 6,000 to 20,000 dollars per kilometre, depending on how many fibre pairs it carries and how heavily it is armoured. Keep in mind that the spending does not stop once the cable is lit, because a submarine cable has a design life of only around 20 to 25 years and on top of that there are somewhere between 150 and 200 faults occurring across the world’s cables in a typical year. The overwhelming majority of them are not caused by sabotage or sharks, but by the combination of fishing gear and dragged ship anchors. Each break has to be mended by sending out one of a small number of dedicated repair ships, that are on permanent standby under regional maintenance agreements, to grapple the cable up off the seabed, haul both severed ends to the surface, splice them back together, and lower the repaired thing back down. This is slow and weather-dependent work that is quite expensive. With the data provided by TeleGeography ’s Submarine Cable Map I have put together a list of the (co-)owners of undersea cables and sorted it by the number of cables each individual company has a stake in. The full dataset runs to some 473 distinct owners, the overwhelming majority of which are obscure national and regional carriers you will never have heard of, so rather than just dumping the entire list here, I limited it to the hundred most prolific (co-)owners: Note: These figures are derived from the public Submarine Cable Map data, counting both, systems already in service, and those still planned or under construction (603 of the former, 91 of the latter, at the time of writing). The field is free-form text, so a few owners turn up under more than one spelling, and I had to do a little manual untangling of company names. What jumps out, at least to me, is the name sitting right at the top. For most of the history of this infrastructure the owners were telephone companies, the BTs and AT&Ts and NTTs of the world, laying cables to carry one another’s calls and, later, traffic. Google now has a stake in more submarine cables than any traditional carrier on the planet, with Meta not far behind, and Microsoft and Amazon both slowly accumulating their own share. The companies that fill those cables with traffic have, over the past decade or so, decided that they would rather own the pipes than rent them. The other thing the numbers tell you is just how long the tail is. Of those 473 owners, some 260 appear on exactly one cable, and more than 340 of them, north of seventy percent, on no more than two. These are the world’s national telecoms, each one buying a slice of the handful of consortium cables that happen to land on its particular stretch of coast, which is also why so many of the big international systems list a dozen or more co-owners apiece. The internet, seen from this angle, is less of a single network and more of a mix of local operators, all chipping in for a share of the same few very expensive ropes across the ocean. To see what it looks like where the cable actually meets the land, let’s head back to that beach in La Línea . The cable that surfaces there is called Dos Continentes , it belongs to GTD , a Chilean telecoms group , and it’s a relatively small regional system consisting of two armoured fibre cables looping across the Strait of Gibraltar to Ceuta , the Spanish enclave on the African coast that depended on a single ageing link before this one was built. I went looking for exactly where it comes ashore, and the paper trail gives an idea about how invisible this infrastructure actually is. The cable lands in Spain, but the public Spanish government map of coastal concessions doesn’t seem to show it, because it looks like coastal permits in Andalusia are devolved to the regional government. The landfall instead shows in a regional registry , in a signed resolution buried under an expediente number. That document pinpoints where the cable enters the public maritime domain, at grid reference , just seaward of the beach manhole. The cable then runs inland, buried as the permit insists ( “no exterior element above ground level” ) to what is presumably a network node, where traffic is fed into GTD ’s pre-existing terrestrial dark-fibre network, from where it’ll eventually travel to one of the actual GTD data centres in Madrid , Barcelona , Bilbao/Sopelana , and Sevilla . On its way out to sea it crosses three older cables already lying on the seabed, namely Europe India Gateway , ATLAS , and FLAG . As can be seen (or, well, actually not) even an empty-looking patch of water off a Spanish beach is layered with other people’s infrastructure. Note: When GTD applied, it seems that the town council of La Línea formally objected and asked them to drop the project. The cable, the council said, cut straight through the main local fishing ground, “splitting it literally in two” , threatening the small shellfish and trasmallo boats that work those waters, and a protected limpet that lives on the rocks, in a town whose fleet was already squeezed by run-ins with Gibraltar over fishing rights. However, they were overruled and the concession was granted anyway, with mitigation conditions attached, for an initial fifteen years. The Dos Continentes cable ( Segment I , La Línea - Ceuta Sur ramal ), owned by GTD Cableado de Redes Inteligentes, S.L.U. , the Spanish arm of the Chilean GTD group , has a total length of ~105 km and is in service since 2020 under the signed concession resolution from the Junta de Andalucía ( Dirección General de Calidad Ambiental y Cambio Climático ), expediente , dated 14 January 2020. The two key points, as given in the resolution’s coordinate table are: Note: The resolution’s prose text gives a slightly different value that disagrees with its own table by approximately 140m. To convert the UTM coordinates I used the official Instituto Geográfico Nacional ( IGN ) Calculadora Geodésica with the following settings: and differ by only centimetres in practice, so the resulting coordinates (WGS84-equivalent) can be dropped straight into any consumer map or GPS app: Both points sit on Playa de El Burgo–Torrenueva , beside the Punta de Torrenueva tower, at the northeastern ( Levante / Mediterranean-facing) edge of La Línea de la Concepción , against the municipal boundary. The resolution describes the route as passing “muy cerca de la torre-faro existente en la Punta de Torre Nueva” . As you can see, however, you see nothing. :-) The permit requires the whole installation to be subterranean ( “no exterior element above ground level: No manholes, splices, connections or terminals.” ), hence you can stand exactly on the landfall, but it’s a point in the sand by a tower, and not a structure. On the afternoon I was there, a couple of dozen people were spread out on that stretch of sand under parasols, probably not even knowing that somewhere underneath them the link that carries an entire enclave’s traffic to another continent came out of the sea. It is interesting to see that what has changed most over the past decade isn’t the technology itself, but who pays for it. For a century these systems were built by carriers selling capacity to one another, which made the network something close to a shared utility with many owners. Today, however, the largest (co-)owner of submarine cable on the planet is an advertising company. It probably makes sense in their position, however it is a change in how the network is governed, and, more importantly, it seems to have happened almost entirely out of public view, which is worrying. If you live anywhere near a coast, there is a decent chance one of these things lands within driving distance of you, and the TeleGeography map will get you to roughly the right bay. Getting from there to the actual patch of sand takes some amount of digging through concession resolutions, planning registers, environmental reports, and sometimes the local newspaper archive. It took me an evening of reading to narrow it down, but I can recommend to do this exercise if you’re curious about the world that you’re living in and, more importantly, the hidden infrastructure surrounding you. PS: Maybe we picked the wrong word and should have called it the trench rather than the cloud ? Transformation type: Transformación de Datum Reference system: ETRS89 Input coordinates: UTM Huso (zone): 30

0 views

Who’s Tracking You? Use This New Service to Find Out

It can be daunting to determine who’s responsible for showing ads on the websites we visit, or who’s harvesting data from the mobile apps we use every day. That information is already semi-public, but it is not easily parsed and traditionally much of it has remained walled away in the hands of large advertising platforms. Not anymore: A powerful and free new service called DecryptAds scrapes and correlates this adtech data and makes it simple to quickly learn a great deal about the entities that are tracking you. A Decryptads summary of the advertising partnerships declared by espn.com. The newly launched decryptads.com says it is constantly scraping the files that websites and apps make publicly available to disclose the companies that are permitted to run ads or collect user data. These files include: – ads.txt : all of the adtech companies and data brokers that may run ads or harvest data from the site; – app-ads.txt : entities that can harvest data from or display ads on mobile and smart TV apps; – buyers.json/sellers.json : the entities buying, selling or reselling ad inventory for a given site or app. Zach Edwards is chief research officer for DecryptAds and a threat researcher at the security company Infoblox . Edwards said he and two other founders decided the service was needed because the adtech data in these files is generally only useful when it can be cross-referenced to build a more complete picture of the advertising ecosystem for each website or app. “It’s an adtech tool but we’re trying to approach adtech from a security perspective,” Edwards said. “It’s really built for a lot of privacy and security use cases that have been dramatically underserved.” Those use cases, he said, include tracking down the source of malicious ads that try to foist malware on targeted users, identifying ad networks located in adversarial nations, and detecting the fast growing swarms of AI-generated slop websites and apps. And as decryptads.com demonstrates, these potential security and privacy threats are near impossible to detect just by viewing a single apps.txt or app-ads.txt file. “Supply-chain integrity issues rarely live in a single file,” the site explains . “They show up as broken cross-references between ads.txt, app-ads.txt, and sellers.json files; as cloned declaration sets across unrelated domains; as seller removals that only make sense when viewed across exchanges; and even as supply paths in bid logs that never actually appear in any given publisher’s authorized-seller list.” A search in DecryptAds for the hugely popular sports network espn.com reveals 143 ad partners and 19 registered data broker domains are listed within its ads.txt and app-ads.txt files. That data broker information is gradually becoming available because four states — California, Oregon, Texas and Vermont — have recently passed laws requiring data brokers to register if they buy or sell data on consumers from those states. DecryptAds reports that almost half of those data brokers are collecting geolocation data from espn.com visitors who aren’t blocking ads, while another three disclose that they collect device fingerprints and sensitive personal information. A visual representation of the complex ad supply chain declared by espn.com. Image: decryptads.com. DecryptAds also makes it easy to learn the beneficiaries and national origins of the advertising firms lurking in apps and websites, displaying a conspicuous warning when adtech partners of an app or website are based in “geo-risk” areas like China and Russia, or in countries with strong financial and political ties to both — such as Cyprus and the United Arab Emirates (UAE). According to DecryptAds, espn.com works with four different advertising entities that are based in either Russia, China or the UAE, including the adtech firm Between Digital , which lists a New York address. However, the dossier on Between Digital flags them as a Russian firm, showing that their publisher offers (PDF) are processed through Alfa Bank , Russia’s largest private commercial bank and one of several financial institutions placed under U.S. sanctions in 2022 after Russia invaded Ukraine. KrebsOnSecurity sought comment from both Between Digital and the company’s founder, and will update this story in the event that either replies. A search for several top U.S. military news websites — including armytimes.com , airforcetimes.com , defensenews.com , navytimes.com , marinecorpstimes.com and federaltimes.com — shows they all allow Between Digital to serve ads and track users, as well as two entities in the UAE and another in the ownership secrecy haven of Panama. DecryptAds reports that Between Digital is collecting ad data on approximately 55,000 partner websites. The “Geo Risk” section of decryptads.com. Pivoting on Between Digital’s app-ads.txt file reveals hundreds of domains featuring simple web-based games that are frequently interrupted by ads. Edwards said Between Digital’s own declarations show the company is listed as both a publisher and a reseller on approximately two-thirds of their portfolio. “It means they are basically playing both sides of the bidding equation, which creates opportunities to direct client spend at your owned and operated properties or client infrastructure, essentially creating opportunities for conflicts of interest,” Edwards told KrebsOnSecurity. “The problem we have right now is that for years we’ve had almost no one policing these ads.txt and app-ads.txt files.” The Opera Web browser remains quite popular, and probably many users are unaware that since 2016 it has been majority owned and controlled by the Chinese company Kunlun Tech (the operational headquarters of Opera remain in Oslo, Norway). Opera.com’s profile at DecryptAds identifies 27 registered data brokers collecting information, including 15 adtech partners in the UAE, six in China, three in Cyprus, two in Russia and one each in Hong Kong and Ukraine. DecryptAds makes clear, however, that these companies represent just seven percent of the adtech partners specified in Opera.com’s ads.txt and app-ads.txt files. One feature of DecryptAds that sent this author down multiple hours-long research rabbit holes is its Legal Dossier lookup , which takes several minutes for each search but eventually churns out oodles of useful information about who owns a particular domain or app, when it was registered, and any aliases or relationships it may have to adtech companies and other websites or apps. For example, last month KrebsOnSecurity wrote about researchers from Bitsight who found that an extremely popular line of TV streaming sticks called H96 quietly rent out each user’s Internet connection to strangers. Bitsight also discovered that when these devices aren’t being used to stream pirated video content, they are spoofing themselves as mobile phones clicking ads on AI-generated slop websites . Bitsight concluded that the same Chinese company that made several of the malicious apps common to all of these H96 streaming sticks — the Fengwo Group — also also ran the network of ads and AI slop websites being clicked on by tens of thousands of these devices that are pretending to be mobile phones. Examples of ad landing pages linked to the Fengwo Group. These sites were designed to show ads only to H96 devices that were spoofing their device type as mobile phones. Image: Bitsight. A DecryptAds legal dossier on the (now dormant) Fengwo Group domain name for the AI slop website pictured on the left in the screenshot above ( medicalbeautyhub dot com ) shows it shares a seller ID ( 1674071 ) with a gaming website — giacoloredstones[.]com — which features yet another seller ID ( 103488000 ). Pivoting on that latter seller ID reveals hundreds of active websites within Russia’s Yandex ad system featuring extremely low-quality games or simple utilities that pepper visitors with ads. Edwards said that when advertising networks suspect a given advertiser is engaged in unauthentic clicks or displaying malicious ads, very often those networks will quietly remove the offender from their list of approved partners without letting anyone else know about their suspicions. This practice, he said, makes it easier for dodgy adtech firms to avoid accountability and continue victimizing others. To address that visibility gap, DecryptAds features a quiet removals feed that records and correlates all of the sellers.json removals across ad exchanges for the same seller domain or name. A screenshot of the Quiet Removals Feed at decryptads.com. “The way the adtech industry works, someone will write a report about ad fraud and only share it with their own clients and they won’t make it public,” Edwards said. “The ban is just removing them from the sellers.json file, but they told nobody. One day it was there, the next it was gone. So if you’re trying to navigate who is suspicious, that’s usually tough to do because there are a lot of adtech companies removing things all at once.” Malvertising, the term given to the practice of inserting malicious ads that foist malware or redirect visitors to phishing pages, remains an all-too-frequent occurrence in the modern adtech industry. But Edwards said these malicious ads are far more commonly found now on newly generated AI slop websites than on high traffic destinations that typically employ a variety of technologies and third party tools to quickly flag bad ads. “None of these slop AI content farms are paying for that kind of protection,” he said. “They’re just signing up the lowest quality partners, and it essentially becomes a greased rail to target the users of those sites with malicious ads. Most malvertising attacks don’t happen on espn.com or huffpost.com, but rather [on] some lower quality content farm and someone just went there because it came up in a search.” Edwards said the AI slop websites are populated with machine-generated blog posts and images, and cover a wide array of themes from home improvement and decorating to food recipes, hunting, cars and consumer technology. He said organizations that get hit with malicious ads are often at a loss for what to do next, unaware that in most cases the answer is one of the entities listed inside the website’s ads.txt or app-ads.txt file. “A lot of serious organizations are starting to understand that if we’re not breaking down this ad data, we’re not going to know who’s targeting government people with zero-click payloads on an almost daily basis,” he said. Edwards maintains that truly getting a handle on the malvertising and AI slop problems will require more data-sharing by the major ad networks. Specifically, he says those platforms do not broadly share what’s known as the “supply chain object” or SCO, structured data attached to each advertising bid request that lets buyers see every seller, reseller and intermediary involved in passing an ad impression from the publisher to the final buyer. “That SCO tells you who sold it or resold it, and who was the final entity that bought the impression that served that malware payload,” Edwards explained. “You may see the malicious zero-click redirection, but without the supply chain object — which is only served server side — you won’t know who targeted your people with malware and won’t have a way to try and prevent it properly. But if we can encourage the adtech industry to expose that SCO, it will get easier to find the culprit behind any one bad ad.” DecryptAds also offers an application programming interface (API) that allows researchers to automate queries and integrate the site’s functionality into popular AI platforms. The only sane reaction to the examples described above is to block all online ads outright. This approach is broadly endorsed by security experts because it also makes it more difficult for adtech firms and data brokers to build detailed profiles on you and track your movements around the web and in the real world. However, much depends on how you normally prefer to browse the Internet, and how much trust you place in third party browser plugins and extensions. For those primarily surfing via a regular desktop or laptop Web browser, uBlock Origin Lite is an excellent free and well-maintained open source option. uBlock Origin also should work with mobile browsers like Firefox, but apparently only on Android-based devices. Adblock Plus is a decent option for iPhone and iPad users. For power users, Adblock and uBlock Origin both support custom blocking rules from easylist.to , which publishes a frequently updated list that removes most advertisements from webpages. The well established browser extension NoScript blocks all non-approved Javascript code, and it generally does a fine job blocking most ads from loading. However, script blockers like NoScript may not be suitable for average users who don’t enjoy constantly having to referee which scripts should be allowed to load so that each site displays properly. More technically inclined/adventuresome readers should strongly consider a hardware approach to blocking ads at the local network level, because that is easily the cheapest, most secure and scalable way to do it. A tiny, low-cost and broadly available computer known as a Raspberry Pi can be turned into a powerful ad blocker for all devices on a local network when fitted with a microSD memory card and a free program called Pi-hole . Once you’ve set it up properly and changed your router’s network settings to use the Pi-hole’s DNS sinkhole and DHCP servers, it should prevent ads from displaying on any devices connected to that network. Bear in mind that ad blockers often do little to block ads and/or tracking that occurs from within mobile apps that users have chosen to install on their devices. Many websites now push users to install a mobile app, supposedly in order to more fully access and enjoy the site’s services and content. But in my experience, they’re not doing this because the user experience is somehow way better on the app (as LinkedIn tries to convince us non-app users several times a week via email). On the contrary, I find most mobile apps to be horribly designed, annoying, and/or completely unnecessary, and when given the option I will almost always choose to interact with a website or service directly in a Web browser. No, the cold truth is that big web destinations tend to get pushy with their apps because they make it easier for these companies to keep you on their platforms longer and to collect (and in many cases resell) far more precise data about who, what and where their users are. Also, companies pushing customers the hardest to install mobile apps always seem to liberally opt everyone in to having their data used to train large language models these days. So be cautious about the apps you install on your mobile devices ( including any smart TVs! ), and poke around their listings at DecryptAds if you want to learn more about their privacy practices and any relationships they may have to adtech firms.

0 views
iDiallo 2 weeks ago

Where Did the Productivity Gains Go?

There is a machine called Productivity, and on this machine there is a knob and a button. The knob increases productivity, and the button is labeled "reset." My job was data entry at a non-profit. It was my first job in the US where I didn't have to lift heavy objects. My main task was entering collected donations into a database. The data source was email attachments, which came in various formats. Excel files, Word files, and some directly embedded in the email body. The database itself was FileMaker Pro. Every part of the process was wasteful. To receive the emails in the first place, I had to call the person who collected the donation, argue with them, and guide them through the process of emailing me the information. The donations were collected at events, and the data was often just handwritten on a note. Once I got the email, I would extract the content from the attachment and organize it into an Excel file. Then I would manually enter it into the FileMaker database, one row at a time. In a full workday, I could process no more than a dozen entries. So I decided to improve the process. I started from the end. Entering data into the database was error-prone. I had to enter the name, address, donation amount, and a plethora of other information into a single row of a database viewer, and I often put the wrong data in the wrong column. So instead, I created a form with validation. I built a standardized Excel file where the people collecting donations could enter data directly into the correct columns and send it to me. This turned out to be too much to ask. Several people didn't know how to use Excel, and it became a nightmare of technical support. Eventually, I switched to a formatted Word document. Annoying, but it worked. Lastly, I created a process. Every Monday, I would email everyone to remind them to send me their donations by Wednesday, so I would have plenty of time to enter them on schedule. The deadline was arbitrary, but it slowly hardened into policy. My system was a success. I was entering hundreds of donations every day. In terms of productivity, I had more than 10x'd my output, which I thought was a good thing, until I started getting noticed. What have you done for me lately? My productivity machine's knob had been dialed to 11. But I had also set the new standard for how much needed to get done in a day. So my manager hit the reset button. The dial went back to zero, but the expectation stayed exactly where it was. It's funny how that always happens. Increased output never turns into more free time. Instead, it becomes the new normal. Anything less looks like a drop in productivity. These days, with AI common in the workplace, every semi-technical manager has started building apps themselves instead of taking requests to the dev team. They get a rush from how much code they can produce, and they enjoy that high for a while. But they fail to see that the productivity gain only exists at the very beginning, before the app is actually in use. Once people start using it, you don't get to regenerate the app from scratch every time you want to add a feature. Instead, you have to add things slowly and carefully, without breaking what already works. All fields start green, but they brown eventually. When you get an extremely productive teammate, you start looking at the rest of the team as time-wasters. In fact, once you get used to what that teammate produces, you start treating it as the baseline and eventually you ask, "What have you done for me lately?" I was let go from that job in a little celebration, right after I'd trained my replacement. One morning, my manager simply announced that it had only ever been a summer job, and summer was ending. So my last day was set. I was too green to point out that this had never come up during hiring, and that I'd actually started well before summer began. They brought cake on my last day. What they failed to see was that my process was still very much manual. My replacement was let go when she couldn't manage more than a dozen entries a day. So much for productivity gains.

0 views
Tara's Website 1 months ago

Cobolito/400: a tiny data appliance in the IBM 5280 lineage

Cobolito/400: a tiny data appliance in the IBM 5280 lineage I wrote an essay about Cobolito/400, a tiny data appliance built as a tribute to the IBM 5280 and to a style of computing where records, screens, media, and operators still had visible relationships with each other. At the beginning, it was meant to be a blog post. But somewhere between Scottish biscuits and a cup of tea, I got carried away and wrote almost 90 pages instead.

0 views
Stratechery 1 months ago

Muse Image, Grok 4.5, Alex Karp on CNBC

The batter for verifiable data is increasingly defining the AI race, from Meta to Grok to the frontier labs.

0 views
iDiallo 1 months ago

Amazon Basics, but for intellectual property.

Amazon has been accused several times for ripping off merchants on its platform. And every single time they denied any wrongdoing. A merchant, or anyone really, can create a product (or source it from China), then resell it on amazon. Amazon is the service provider, and hosts all the metrics concerning the products. If Amazon themselves were in the business of creating and selling products, then that creates a potential of conflict of interest. Because they have the data of all products that sell and sell well. They could replicate that success without doing any further research since the merchant has already confirmed the existence of demand. It's not surprising that Amazon Basics quickly became the best selling "private-label brand" on Amazon. They already know what sells because they have access to the data. Yet they continued to deny it, and state that they only ever use publicly available data from sellers . An Amazon spokesperson said the company believes the allegations are "factually incorrect and unsubstantiated," adding that Amazon strictly prohibits the "use or sharing of non-public, seller-specific data for the benefit of any seller, including sellers of private brands." Yet the results are right there for all to see . If you sell any product through Amazon, you are exposing your company's operations to them. If you want to keep that information to yourself, then you don't get to reach your customers, which in reality are Amazon's customers . If you want to buy something online, and get it shipped as quickly as possible, then Amazon is a blessing. Most often than not, you are not buying the product directly from Amazon. An independent store or vendor with a presence on Amazon will fulfill your order. The seller only has minor identifying characteristics on the platform. On the search result page, the space designated to the seller is small and insignificant. The customer has very few reminders that products are offered by anyone but Amazon. (Although if you want to dispute a sale, you are starkly reminded that the item is from a 3rd party vendor.) So there is no surprise when companies embrace AI internally, they are putting themselves at the risk of sharing their product with their competitors. Maybe the most obvious example is when Antropic came up with Claude Design. A tool to help users generate designs, wireframes, etc. Kinda like Figma. That's not a problem on its own, but when Antropic's chief product officer sits on Figma's board of directors, you can't say that there isn't a conflict of interest there. In fact, the chief product officer resigned from the board merely days before Claude Design was announced. He basically extracted all value from Figma then resigned. Figma's AI features are built on top of Claude. So Anthropic literally pulled an Amazon Basic on Figma . When companies force their own employees to use AI to do their day to day work, they are basically asking employees to upload company data to a 3rd party that may become a competitor. Sure something in the contract clause says that the AI company won't train on enterprise customer data, but nothing stops them from peaking at successful product data. Whenever someone tells me that they used AI to build an app and boast of its values or uniqueness, I want to remind them that if you can just prompt-create a product, so can the AI provider. In fact, they might have better resources to create a competing product if it displays any sign of success (see Figma). While it looks like plenty of people are benefitting from AI today, all this information is being shared with AI providers. We are giving them full access to our thought process. When you include them in your workflow, you are basically providing them with a step by step approach on how to do your job. Don’t be surprised when you see a native Antropic/OpenAi project management application suite. Or a CRM, or any software that is trying to integrate with AI and may experience success. A few years back, when I worked in Customer Service Automation, we discovered that most companies used Zendesk to manage their customer service. Since customers mainly contacted support via email, an intentional database had been built that tracked users through their shopping experience throughout the web. While so much could be done with that data, like identifying “problematic” customers, or recommending products based on their history, we ended up finding something more helpful. We could easily detect a pattern of issues for certain shipping carriers. We could see when UPS was having delays in certain cities, or when Fedex was having technical issues when updating the last mile status. None of these things were features designed or provided by anyone. However, having access to businesses’ data gave us insight where we had none before. That became a feature for us, only because we were not competitors to all these online retailers. When you expose your company's internal data to a potential competitor, don’t be surprised when they build a competing business to rival you.

0 views
The Jolly Teapot 2 months ago

Unfinished, part deux

Two years ago, I published a post entitled Unfinished . It was a way for me to share some thoughts without having to work on them as much as I do on regular posts. As I wasn’t sure if these “lesser thoughts” were worth my efforts and my time, I compiled them in a different post format, inspired by a song: This post is inspired by the excellent track entitled Lamb’s Garbage (Unfinished) , from the classic album and one of my favourites, Mr Oizo’s Lambs Anger . The concept of the song, as its title suggests, is to regroup bits of songs that were never completed to be full tracks. Well, here we are again. The text file where I jot down all my ideas, quick thoughts, and potential topics for blog articles is starting to get a bit too long for my liking, so I think it’s time for a little clean-up. What you will see below is what was saved from the big flush, and what I don’t share on social media since I am no longer participating . Think of this as a list of intros, tweets, and blurbs of what was going on in my head recently. I believe some of these themes can be used later for a full post; in the meantime, feel free to use them for your own blog. And if you don’t have a blog, please, start a blog . I love spreadsheets. This is something I find a bit difficult to admit, but I do like working in spreadsheets. I even firmly believe that Google Sheets is their best product. I already like lists, but a spreadsheet is on another level. I like to make my spreadsheets look pretty, I like to plan how they will look, I like to build, I like to make them functional, legible, easy to read. For me, it’s a very pleasing and interesting thing to do at work: there are so many possibilities. When I create a spreadsheet, I feel like an app developer. I feel like I’m a graphic designer. I had a co-worker once whose job included the creation and design of very complex spreadsheets for other teams, using Microsoft Excel, Microsoft Power BI, and such. The resulting spreadsheets were glorious: fully featured and interactive dashboards, gathering data from different sources in real-time. Works of art. Are answers from A.I. chatbots recycled for other users asking the exact same thing, or are answers always generated from scratch? Wouldn’t it be cheaper and more energy-efficient ? If I ask “ explain the difference between irony and happenstance ”, will the A.I. chatbot just paste an existing, perfectly fine answer (one that received positive feedback in previous chats), or will it work to generate a brand new answer? Why do so many people keep saying “Samsung charger” or “iPhone charger” instead of USB-C, USB Type C, or just USB? I mean, despite these cables and connectors being ubiquitous in our lives, I see a lot of people completely ignoring what they are called. I wonder why. Don't brag so much about using A.I. It’s great that you used A.I. to do this thing you’re presenting. I can see how it has been useful and how much faster it helped you reach your goals. I understand that without A.I. you could never have pulled this off. I know it’s a way to show how you are part of the A.I. revolution, that you’re not left behind. No shame in that. I work with A.I. a lot too, I’m not judging you for that. But please, don’t present your use of A.I. as a skill. It’s just a tool. Your skills are elsewhere. Having access to tokens is a weird flex. The tools you use and how you use them may interest a few of your peers, but what you create with these tools is what truly matters. Do you know what type of video cameras were used in your favourite film? Do you care? By the way, the same piece of advice applies to air fryers. If we work so hard on automating our current tasks and projects with A.I. agents, how will we tell which ones are worth doing at all? Does everything need to be A.I.-enabled and optimised? Are we reproducing the same mistake that we made with social media, shoving it everywhere we could? On that topic, I highly recommend this excellent article on The Verge . Efficiency is not the ultimate goal for most people: efficiency for what? For whom? Besides, friction is not always a problem : sometimes friction is how new ideas spark to life. If you are like me, an avid consumer of Techmeme , you will have noticed that A.I. companies get a huge part of the coverage these days. I don’t know if it’s an editorial choice of Techmeme or if it’s just a reflection of the public reception of said news, but my gosh it seems that Gemini or ChatGPT or Claude gets an incremental update every day, and they float on top of the site’s homepage seemingly forever. I wouldn’t mind a new site just for A.I. news, just like Mediagazer does what Techmeme does but for everything media-related. I’d call it Datacenter and it would make Techmeme a bit more interesting. I recently discovered that something I immensely dislike has a name: the Rae Dunn style for household items. Billionaires cannot stand the idea of a democracy where their individual vote is, technically, worth exactly as much as the vote from the person who takes care of their laundry. They hate that. So what do they do? They buy media or social media companies to try to influence thousands to vote like them. Side note on the ridiculous LinkedIn habit that consists of putting a link in the comments of a post, and writing in the post “Link in the comments”. Just put the link in the post, as you’re supposed to, so we can have a nice preview of the post, and we don’t have to look at the even more ridiculous comments of every LinkedIn post. How messed up is that? I know it’s for better “reach” and to trick the algorithm, but you just look thirsty for likes. Isn’t that link the thing you wanted to share? Do you prefer a click or a like? What’s a like good for if nobody visits your link? Thankfully, I don’t have a LinkedIn account, and I can ignore this nonsense most of the time, but I do check on a few LinkedIn posts for work and this is making me both sad and angry. On Instagram, the whole “Link in bio” was necessary because that was the only way to share links back then. But LinkedIn? No excuse. Yes, it sucks that their algorithm prefers posts that won’t send users out of their precious, shitty platform. I’m with you. But you don’t have to play their silly little game. You’re better than this.

0 views
Aran Wilkinson 2 months ago

Headcode started as a project to learn UK rail data

It now has eight live endpoints , a tiered pricing page and an alpha banner warning people the schemas might still move under them. None of that was the plan. The plan was to understand how UK rail data actually fits together, and to work out how to use an LLM properly on something real instead of a toy. The product is just what happened while I was doing that. I'm still not sure it becomes a business. I'm completely sure the way I learned to work on it was worth the time, and that part I now use every day on everything else. One station has more names than you'd believe, and that's the whole problem. King's Cross is KGX to the fares system, KNGX to the timetable and 54311 to the movement feeds, with a handful more codes besides (NLC, ATCO, UIC), each from a different corner of the railway. You don't need to hold those in your head. Nobody ever agreed on one name, so every feed brought its own. It gets worse with size: a big terminus like London Bridge doesn't have one timetable code, it has a cluster of them, roughly one per platform group, so even "which code is the station" isn't a clean question. Then there's the live side. Darwin (the real-time running feed) pushes updates at you as a stream, while the reference data turns up separately, each source on its own schedule and in its own shape. Before you can render a single departure board you've written a reconciliation layer and a code-mapping table, and now you own both of them. That mess is exactly why it was good to learn on. Bounded enough that you can actually finish it, awkward enough that you can't bluff your way through. You either understand how the feeds relate or your departure board quietly shows the wrong train. The workflow I use now didn't arrive fully formed. It evolved, and the project is where each step earned its place. I started where most people start: a prompt and a plan. Ask for a thing, get a plan back, let it build. That's fine for small, self-contained work. It fell apart the moment a task touched code the model hadn't really looked at. It would produce something plausible and confidently wrong, and then I'd spend longer unpicking it than the thing was worth, either fixing it by hand or trying to prompt my way back out, which sometimes just dug the hole deeper. The failure wasn't the model being bad. It was me asking it to act on intent it didn't have. So the requirement moved to the front. Before any code, I'd work up a proper PRD with the model: I'd set the direction and push back, it would draft and fill in. These weren't a paragraph of good intentions. They grew into real documents, with the goals stated plainly and the out-of-scope list stated just as plainly, functional requirements (what it does) sitting next to non-functional ones (how fast, how reliable, what the limits are), a sketch of the technical architecture, the phases it would be built in, how it would be tested and what might go wrong along the way. Then I'd break that down into tasks small enough to review one at a time. The output got noticeably better, because the model was working to a brief instead of filling in the gaps itself. It stayed focused on the thing in front of it, and it had helped write the thing that kept it there. The step that changed the most came later, and it wasn't obvious to me at the start. A PRD is only as good as your understanding of the code it lands in. So before writing the PRD, I'd have the model research the existing codebase and write up how the relevant part actually works, what's already there, which patterns to follow. The thing that made the research useful was a hard rule: document the codebase as it exists today, and nothing else. No suggested improvements, no root cause analysis, no critique, no refactoring proposals, no architecture it wished were there. Only what exists, where it lives, how it works and how the parts connect. Left to its own instincts a model will reach for the fix, because pointing out problems reads as helpful. Forbidding all of that kept the output objective, a technical map of the system as built with the opinion stripped out. Adding that step changed the relationship. The model stopped being the author and became something that I directed, sent to find things out and report back rather than left to decide what to build. The research feeds the PRD, the PRD feeds the tasks, and I'm steering at every handoff. Review runs through all of it. I read the research, the PRD, the tasks, the plans, not just the final diff. That order matters more than it sounds. If the research is wrong, the PRD inherits the error and every task underneath it inherits it again, and by the time you're reviewing code you're three layers downstream of the actual mistake. The review at the top is worth far more than the review at the bottom. That's the honest line between this and vibecoding. It was never about whether I used the AI. It was about whether I ever let go of understanding what it was doing, and I made a point of not letting go. It wasn't all done this way from the start. The early parts of Headcode were built the way most things get built with an LLM, by prompting, getting something working and moving on. The discipline came later. As the research-into-PRD-into-tasks process settled, designing before implementing became the default, and the project quietly split into a scrappy first phase and a deliberate second one. You can see the join. The later work, the bulk of the API surface, the endpoint groups, the schema as it stands now, was designed before any of it was written, the spec settled while the code was still hypothetical. The early prompted bits I've mostly gone back and rebuilt to the same standard, because once you've felt the difference the scrappy version nags at you. What fell out of that patience is an API where the schema is the contract. It's OpenAPI-first, with a downloadable spec you can point a client generator or contract tests at. Identifiers resolve cleanly: hand it any code system and it gives you back all of them, so the reconciliation table that would normally be a thing you maintain becomes a field you read. Vibecoding the same idea would have got me a convincing departure board demo and a wall the moment the identifier resolution got hard. Rail data punishes building before thinking, which is precisely what made it a good teacher. Headcode is in alpha. There's no self-serve signup. Access is by request, you email me with what you're building and I send a token. That's deliberate while the data and the schemas are still settling. I genuinely don't know whether there's external demand for it. It might just stay a personal project, something I experiment with and build other things on top of, now that I've got the rail data in a clean format to start from. That alone was worth doing. I want to build visualisations on top of it, possibly a small app, and a clean API I control is reason enough to have built the thing. If people turn up actually wanting the data, it might grow into a small SaaS. What I'm not going to do is manufacture a roadmap I don't believe in, or pretend there's urgency around a project I started in order to learn. Whatever Headcode turns into, the workflow has already paid for itself. I went in wanting to learn how to use an LLM well on a real codebase, and the research into PRD into tasks pipeline, with review at every layer, is now simply how I work with one. The product is a maybe. The method, I kept. If you want to see what came out of it, Headcode lives at headcode.dev , and the API docs — endpoints, schemas and the OpenAPI spec — are open to browse at docs.headcode.dev .

0 views
Jack Vanlightly 2 months ago

Can We Agree on a Storage/Workload Architecture Taxonomy?

The lines between transactional systems, analytical systems, hybrid systems, and shared storage architectures are getting blurry. This post proposes a small taxonomy for describing the different ways systems, workloads, storage tiers, visibility, and durable copies relate to each other. OLTP, OLAP, HTAP, and now LTAP? We can think of the first two as two types of workload which have specialized query engines and storage systems to support them. OLTP such as the RDBMS like Postgres and MySQL use row-based storage engines. OLAP, such as Clickhouse, cloud data warehouse and the lakehouse use column-based storage. HTAP is a hybrid workload system: one system -> both transactional and analytical workloads. The HTAP system therefore has specialized storage and specialized query engine to stitch together the row-based and columnar data. So far, we’re dealing with a single system. A Postgres (OLTP), a Clickhouse (OLAP), a SingleStore or TiDB (HTAP). So what is the recent Databricks’ LTAP announcement? LTAP is the two workloads (OLTP and OLAP) but also two systems (e.g. Postgres and lakehouse/Spark) and some blend of two different storage systems. As well single single vs multi-system, single vs multi-workload, there are other relevant concepts such as tiering and materialization: A single system can tier (move) data from hot to cold storage (for cost efficiency). One system, one copy, two tiers. Hot and cold might be the same storage format (both row-based or both columnar), or might be different formats (hot is row-based, cold is columnar). We can have two systems share the same storage tier. System A tiers (move) hot data to the storage of System B. Two systems, one copy, though System B doesn’t see the newest data yet which only exists on A. Materializing One system can materialize (copy) data into another system. Two systems, two copies. Note when I say “copy of the data”, I mean durable copy, so caching doesn’t count. If the number of copies really matters to you as a metric, then maybe caching does count, depending on how much cached data you need to make it work? If only life were simpler. It would be nice to have some shared vocabulary around this, so we can talk about system architecture more easily. So I defined some terms last year for this, and expanded it as seen below. Vis means Visibility (when is data available in the other workload). The broad classification scheme: Single tier, one system, one workload. Example: Postgres with SSD, single tier CockroachDB, standard Kafka cluster. Internal Tiering, one system, one workload, commonly tiers from hot to cold storage for cost efficiency, e.g. hot=SSD, cold=S3. Though tiering could also serve other purposes than cost. Example: Apache Kafka tiered storage, ClickHouse MergeTree tiered storage. Hybrid-Sync (aka HTAP), one system, two workloads, two or more storage with potentially different formats/tiers, e.g. hot row-based data on SSD, long-term columnar data on S3. Data is immediately available to both workloads (e.g. OLTP queries and OLAP queries). Example: SingleStore and TiDB (Pingcap). Hybrid-Async , one system, two workloads. Like Hybrid-Sync except hot row-based data is asynchronously tiered to long-term columnar format. OLAP queries do not see the very newest data. Example: Snowflake Hybrid tables. Materializing , two workloads, two systems, two copies. System A copies data to System B. Each system is dedicated to one workload, with specialized query engine and storage. Example: ETL in general, many Kafka-compatible services have automatic Iceberg materialization of topics e.g. Confluent Tableflow, Databricks Synced tables asynchronously materialize from lakehouse to lakebase (Postgres). Shared Tiering , two workloads, two systems. one copy across hot tier + shared colder tier (e.g. hot row-based data on SSD for System A, colder columnar data on S3 for System A + B). Example: Apache Fluss tiers hot data (Fluss servers) to lakehouse (lakehouse is a shared tier), LTAP. Potentially, a 7th and 8th category could hypothetically exist: Shared-Sync-RR and Shared-Sync-MM. Two systems, two workloads, one synchronous storage (each write is immediately visible in the other system. Read-replica (RR) variant has one master system and one read-only system (e.g. writes to Postgres are immediately visible for reads in lakehouse). Multi-master (MM) allows both systems to write (hard!!). At the time of writing the details on LTAP are scarce, but it seems like LTAP will fall into Shared Tiering. The thing that differentiates HTAP from LTAP is that HTAP is a single hybrid system which makes data visible to both transactional and analytical queries at the same time. LTAP is a way of unifying the data of two different systems (each targeting a different workload) and sharing the colder data such that there is no (durable) data copy required. It is fundamentally asynchronous: hottest data is only in System A and the remaining colder data is stored in System B but made available to System A (as it’s cold tier). Of course LTAP could potentially move towards the hypothetical category Shared-Sync-RR , given both systems exist in the same platform, then it gets murky again because its one platform, its veering towards HTAP (Hybrid-Sync). One thing that the marketing material of unified OLTP-OLAP system commonly glosses over are the different data models used in each, such as Third Normal Form (3NF) common in OLTP and Kimball (star and snowflake schema) common in analytics. This adds another dimension, on top of query engine, storage layout and storage substrate. If you want 3NF for OLTP and Kimball for analytics, then it’s probably going to be Materialization (as star schema is not viable as a cold tier for 3NF). What you you think of this broad classification scheme? Find on me social media :) ps, some thoughts on data copies… With Shared Tiering, you can think of the data-copy question as a dial: Dial it to no-copies-at-all means evicting data as soon as it has been tiered. Lower storage cost, but maybe it would be good to hang onto to the hot data a little longer for performance. Dial it to lots-of-data-overlap means aggressively tiering to System B but hanging onto the data in System A for the better performance profile, at the additional storage cost. And technically it would now count as cached data which might not count as a data copy, depending on how you define that. However, the data-copy question is also murky with Materialization. Because we have two (or more) independent systems, each can potentially use independent data expiration policies. For example, in Kafka, it might store 7 days, but in the lakehouse, it might store 7 years. In that case, while theoretically it is a two-copy system, the total duplication would only be 0.0027%. I generally dislike the whole “zero-copy” or “one-copy” thing, it’s too much marketing. Focusing on how many copies you have is just weird as a primary design point when you’re building data systems, the real world is more nuanced. Tiering A single system can tier (move) data from hot to cold storage (for cost efficiency). One system, one copy, two tiers. Hot and cold might be the same storage format (both row-based or both columnar), or might be different formats (hot is row-based, cold is columnar). We can have two systems share the same storage tier. System A tiers (move) hot data to the storage of System B. Two systems, one copy, though System B doesn’t see the newest data yet which only exists on A. Materializing One system can materialize (copy) data into another system. Two systems, two copies. Single tier, one system, one workload. Example: Postgres with SSD, single tier CockroachDB, standard Kafka cluster. Internal Tiering, one system, one workload, commonly tiers from hot to cold storage for cost efficiency, e.g. hot=SSD, cold=S3. Though tiering could also serve other purposes than cost. Example: Apache Kafka tiered storage, ClickHouse MergeTree tiered storage. Hybrid-Sync (aka HTAP), one system, two workloads, two or more storage with potentially different formats/tiers, e.g. hot row-based data on SSD, long-term columnar data on S3. Data is immediately available to both workloads (e.g. OLTP queries and OLAP queries). Example: SingleStore and TiDB (Pingcap). Hybrid-Async , one system, two workloads. Like Hybrid-Sync except hot row-based data is asynchronously tiered to long-term columnar format. OLAP queries do not see the very newest data. Example: Snowflake Hybrid tables. Materializing , two workloads, two systems, two copies. System A copies data to System B. Each system is dedicated to one workload, with specialized query engine and storage. Example: ETL in general, many Kafka-compatible services have automatic Iceberg materialization of topics e.g. Confluent Tableflow, Databricks Synced tables asynchronously materialize from lakehouse to lakebase (Postgres). Shared Tiering , two workloads, two systems. one copy across hot tier + shared colder tier (e.g. hot row-based data on SSD for System A, colder columnar data on S3 for System A + B). Example: Apache Fluss tiers hot data (Fluss servers) to lakehouse (lakehouse is a shared tier), LTAP. Dial it to no-copies-at-all means evicting data as soon as it has been tiered. Lower storage cost, but maybe it would be good to hang onto to the hot data a little longer for performance. Dial it to lots-of-data-overlap means aggressively tiering to System B but hanging onto the data in System A for the better performance profile, at the additional storage cost. And technically it would now count as cached data which might not count as a data copy, depending on how you define that.

0 views
ava's blog 2 months ago

let's talk about your digital remains!

" When I die, delete my browser history. " ­— Unknown When you die, there are lots of processes in place to deal with your body, your burial, your physical possessions, subscriptions and bank accounts. But what about your digital accounts and possessions? As our lives become more and more digital, taking these into account when tying up the affairs of a dead person is increasingly important. Think about it: This can involve e-mail accounts, social media accounts, messengers, LLM conversations, hard drives, cloud storage, crypto wallets, websites, your digital media licenses, intellectual property you released (like four (F)OSS projects, for example), and more. In a broader sense, you might count browser history and other metadata, too! What's interesting is that so many of these do not fall under the laws you might expect them to, like succession/inheritance law or privacy law. Services that offer you licensed content (like Steam) have made clear in the past that family members are unable to inherit the accounts or licenses, like they would with physical items. In terms of privacy and data protection, the GDPR applies only to living people, so you lose these rights upon death; the task of legislating the rights of the dead in these regards has been given to the Member States, which results in quite a patchwork of rights 1 . This patchwork makes things difficult, because it means your European country can have different laws than another, and companies will have to see how to comply with them all. France, for example, has one of the most developed post-mortem data protection regimes in Europe. The French Data Protection Act (the Loi Informatique et Libertés ) actively considers death in data protection and explicitly allows a person to give instructions regarding the retention, deletion, and communication of their personal data after death and appoint a person responsible for implementing those instructions. It mandates that controllers must follow the deceased's valid instructions, and heirs can obtain access to data necessary to settle the estate, to identify assets and liabilities, or to close user accounts and manage digital affairs. Germany, on the other hand, is pretty much the opposite: Protection of deceased persons' data arises from a combination of post-mortem personality rights (postmortales Persönlichkeitsrecht) concluded from civil law and constitutional law, inheritance/succession law, confidentiality obligations and possibly some sector-specific laws. It's a lot more complicated and full of holes for specific types of digital data. I wish we had a law like France has! Regardless of when we might have a European law harmonizing this aspect across Member States, it's still important to ask yourself: Who is allowed to have access to your accounts and data after you pass? You might still want to give your younger sibling access to your Steam account later, or you need your spouse to be able to log in and keep a personal website up and running, or save pictures from the cloud. For this, you should make sure that the correct people can have access to your accounts in case of death, and know what to do with them. How you do that is up to you: You might set up something that automatically notifies them about how to access your accounts in times of death when you don't check in for a while, or you tell them a physical location where they can find the device passwords and the Master password to your password manager. I personally mention it in my when i die page. Remember to keep this information updated! Some companies and services, like Apple, Google and Meta, offer settings about what should happen after your death (usually called Digital Legacy tools, Inactive Account preferences, or Memorialization). You're able to set a successor/manager, deletion preferences and more, depending on the service. You have to dig a little in the settings, but if you're reading this right now, I encourage you to go find it. Good to know: Despite setting someone as a legacy contact, these companies might still request additional documents to prove that you really died. On the other hand, it's also okay to want things to be deleted, either by family members, or automatically by the platform itself. At CPDP 2026, I participated in a workshop about digital remains, and my discussion partner said that her Instagram feels so personal that it should be deleted upon her death, but something like a LinkedIn she'd keep up. So decide for yourself: What accounts do you want deleted, which ones can remain up/dormant? You should communicate this clearly in a way the people tasked with your digital legacy can see it, and talk to these people about it beforehand, if possible, or set it up in the settings. If you want to keep data up, is there a maximum retention period you want to set so that the data would be deleted afterwards? As a next step, you have to think about the future. The world will move on without you, and even right now as you are reading this, we are building tech that promises to bring people "back to life" via AI. Even just a decade ago, you likely couldn't have foreseen where we are at the moment with tech being trained to impersonate you. So where will we be decades down the line? That may require the restraints you set in a will to be more on the tech-agnostic side instead of just banning very specific processes and products. This is not just about the recent Meta AI thing; there are several companies in this space, as it looks to be a profitable new market niche: Bereavement tech . So, how do you want your data to be processed? Do you want tech to be trained on it? Do you allow the platform, or your relatives, to train AI on your accounts and other data and media they have on you? Your account might keep posting for you as if you were still alive, generate selfies or videos with your likeness, or it will respond to messages people send to it so they can keep chatting with "you". Side note : Does this truly help the grieving process? I guess we'll have to find out. A physical removal of items, telephone numbers going out of business, and a burial help saying goodbye and accept the finality of it all. Yet social media accounts can exist visually unchanged for years afterward, as the platform may nudge you to message them, reminds you of their birthday, or shows memories from a couple years ago out of nowhere on your feed. If we soon have the option to have people posting as if nothing happened to them, they stay stuck how they were when they died forever. If you never have to deal with the deafening silence from the other party, do you ever really have to grapple with death? And will the person die a second time for you when they stop offering the model? Maybe that's something you wanna blog your thoughts about :) It doesn't have to be so personal and focused on social media platforms as well. How about archives? Museums? You might laugh at the idea, but most stuff in museums is by ordinary people; we might not even know their name. Some people become famous after their death and their possessions and likeness are displayed for people to learn about them (for example: Anne Frank). We get great insights from the things they left behind that they thought no one would read, and if we're honest, likely wouldn't have consented to be out there. This will increasingly happen with digital means. How okay are you with a holo-you or virtual avatar greeting people in a museum? You might not care about any of this at all - if you're dead, then what does privacy and the data matter? It's not like it can still affect you! And that's fair. The views on this can be pretty diverse. Others see the digital remains as a digital version/informational "body" that should also be untouchable and remain undisturbed, and that there should be a general right not to "become a bot". Reading papers and studies about this topic is interesting, because it seems if you belong to the current older generations, you are more in favor of deleting it all, while the younger generations want to keep it up 2 . This makes sense: They might have way more online friends they'd wanna keep this up for. Women seem more in favor to deleting everything than men are 2 , which I can totally see; women tend to make a lot of negative experiences online that center the loss of control over their data and misuse of it. Death, without being able to lock down or delete anything based on developments online seems like the biggest loss of control of all. There are country and cultural differences as well. Unfortunately, unless you control the data (your own Mastodon instance, your self-hosted personal website, etc.), you are reliant on these services to heed your/your loved ones' requests about this. As the big social media companies' business model relies on data harvesting and using existing data for new projects and growth, this might be a hard fight in the future, as they see it as their property. Companies can hold the data hostage because of a lack of laws in your region and no goodwill from their side. There have also been cases already where the companies have refused giving access of a deceased's account to the relatives until a court decided they had to. How many would just give up? For good digital hygiene, we should remember death and make it as easy as possible or sensible for the people we leave behind to get the access they need to manage our stuff how we want them to. Organize your data well (maybe you also want to do some recurring digital version of Swedish Death Cleaning ?), leave instructions, set emergency/legacy access when available, include digital assets in your will, decide how your data is allowed to be used after death, especially around AI replicas. Families should talk about this openly, and relatives and nurses should learn to ask affected parties about these things. Previous related entry: plans for your blog after you die Reply via email Published 15 Jun, 2026 France is very invested in this aspect and its data protection authority (CNIL) has made it one of their main points and even wrote a paper on it. ↩ The CNIL paper has some study summaries about this on page 15. Generally speaking, another good study to read is this one . ↩ France is very invested in this aspect and its data protection authority (CNIL) has made it one of their main points and even wrote a paper on it. ↩ The CNIL paper has some study summaries about this on page 15. Generally speaking, another good study to read is this one . ↩

0 views
ava's blog 2 months ago

what i read this week - week 24 2026

As usual - not counting the personal blogs I read :) Not much appealed to me this week. The AI ‘Revolution' is Not a People's Revolution - AI companies overusing the term revolution is just a marketing ploy, and we should challenge it. Banger quote: " Accepting Blair’s revolution requires agreeing that using unconsented data harvested from populations, processed through biased algorithms and presented to people in addictive interfaces that overwhelmingly generate wealth to US elites, is the change that people want. " Trump Signs Previously Shelved AI Executive Order - summary of the EO. Widerstand gegen Kameras - German article about resistance against surveillance cameras; its history, methods and legal consequences. Person in the comments has an interesting tip: A brush, and acrylic paint mixed with sand. UN-Report zu KI-Umweltkosten - German article about how the UN had the chance of holding tech companies accountable in a new report, but instead only asks consumers to adjust their behavior. I am not opposed to also asking people to rethink their consumer decisions (otherwise, I would not resist using animal products, flying, getting a driver's license etc.; if there's no buyer, there's no product), but for the biggest impact, we need to focus on the source and hold companies to a high standard - or ban their business model or product entirely. The report was also seen as low quality by experts in the field(s). Appeals Centre Europe Transparency Report April 2025 - March 2026 - The Appeals Centre is an independent out-of-court dispute settlement body active due to the Digital Services Act; they've only been around for 18 months. If you are in the EU, you can use them to challenge social media platforms’ decisions on groups, pages accounts or other content which has been removed or kept up despite reporting (if it is about anything other than impersonation, hacked accounts, copyright or CSAM, but hopefully those too at some point). Most cases seem to be about account suspensions, nudity, fraud and scams. So far, they have processed more than 24,000 disputes, where 12,000+ of those fell within their scope. The report has some stats about their work, how many times they disagreed with the platform and overturned the result, and more. DSA User Support Guide also by the Appeals Centre; good breakdown of your rights under the Digital Services Act. The platforms are supposed to tell you that orgs like the Appeal Centre exist, but somehow still don't, and many people don't know their rights. Hold them accountable! Know your rights and make use of the newly established bodies. Under Article 20 of the DSA, users must be able to lodge complaints, free-of-charge, against decisions taken by the platforms within the last six months. Dark Patterns in AI Chatbots - self explanatory; basically about design and interaction/output choices that maximize usage and data collection, lie about the capabilities and emotional intimacy etc. I learned a new term: Privacy Zuckering! Also made me read this about Gemini encouraging a guy to kill others, steal a mannequin, and then kill himself. Arbeitspapier Identifizierbarkeit - German BayLfD summary and interpretation around identified and identifiable personal data in edge cases/gray areas, especially around pseudonymous data. What means to identify "count"? Not just your own! There's a difference between relative/subjective identification and absolute/objective identification. Sidenote: Love that they recommend RSS-Feeds or a Mastodon Account to keep up to date on legislation in this. From intent to action: the leaders' guide to building AI-powered workplace - paper sponsored by Kyocera and done by Economist Impact, based on a survey of 639 senior executives conducted in October and November 2025, with in-depth interviews with businesses and "thought leaders" in AI, digital transformation and workforce strategy. So... take it with a grain of salt, it is very corporate and very incentivized to be pro-AI in the workplace. Their key findings show that they want more investment, more adoption. But: Despite the "propaganda" (so to say), it exposes a lot of weaknesses everyone is already talking about in the workplace. To name one thing I scoffed at: Page 12, the fact that so many measure ROI of AI use in vague "employee productivity", which is probably just increased output or increased closed cases, without looking at the quality. Sad. 4% are not even measuring any ROI for it! Our Data After Us - paper by the CNIL about our digital remains. Covers questions like: Do you want the content to remain after your death? Who gets to have access and manage it? Should that person delete it, or should the platform automatically delete it? Should your remains be used to train an AI to impersonate you to help your loved ones? There seem to be age and gender differences to these answers. You Trust Your Chatbot With Everything - Should You? - paper by Theodore Christakis from AI-Regulation.com. The findings are as expected: Every major provider now trains on consumer chats by default, providers typically reserve safety and abuse-prevention uses and feedback actions to override the training opt-out, and they all reserve the right for humans to read the conversation. The author suggests a " Sealed Mode " where the default settings/options constrain reuse and human access, allows no training, has no advertising, little personalization, and cryptographic hardening. In my view, it could be a good first step, but I fear in practice, it would be bastardized, as meaningless and misleading as Incognito Mode in browsers has been. Ideally, the things of a Sealed Mode should be the default you can then opt out of one by one, and it can be legislated so. We have seen that hidden settings within different menus and specific modes you have to first know about and then turn on do not help the average user, since they are never actively prompted about them or told about them by the company. This stuff only aids a risk-transfer from controller to data subject. So do not offer a silly little compromise - make them default, and do not allow it to cost anything. Choosing between payment or privacy sucks. We should sometimes ask ourselves: If LLMs are just another tool, would I want Microsoft to always have access to and review my Excel sheets? Of course not! So why should we accept this here? At times, the author is too timid for me (" Yet the purpose of adopting this prism is not to export the GDPR as a universal template, nor to argue that the world should converge on European legal categories of individual control. " hey, why not? We don't have the Brussels Effect for nothing; privacy legislation worldwide has been shaped by the GDPR, one example being Brazil!). Favorite chapter was the second one (Ghost in the Machine), as it goes in on how incomplete and lacking the warning labels are, together with how contradicting they are when everything else encourages you to freely share anything. Least favorite are the parts where chatbots are asked to answer something; I am sorry, but I will never see these as genuine, truthful, verifiable answers. This is treating them as a conscious employee that an regurgitate internal policies, not a probability machine who can be nudged to give specific output. Gewalteskalation als System: Nihilistic Violent Extremism in Deutschland - German paper on NVE that's mostly done by children and teens, who connect online over misanthropic and nihilistic tendencies and then see extreme violence and vandalism as the only way forward. Not always far-right or incels, but often. The paper explicitly mentions the Com network, 764, MKY and NLM. Aside from Telegram, Discord is the biggest place for it. I was surprised how lax and wide the definition of violent extremism is (imo, that would make a significant portion of the population violent extremists), and I think the way the authors narrow it down a bit is a good attempt. 28.05.2026 – 26 O 869/26 aka the big one currently making the rounds about Google being responsible for the AI summary output. It will be interesting to see how that progresses and if it will be overruled. This one for noyb. In total, that is roughly ~ 350 pages, if we count an online article as two pages on average; difficult to judge for 17776, I'd put it as 40 pages, maybe. Reply via email Published 14 Jun, 2026 The AI ‘Revolution' is Not a People's Revolution - AI companies overusing the term revolution is just a marketing ploy, and we should challenge it. Banger quote: " Accepting Blair’s revolution requires agreeing that using unconsented data harvested from populations, processed through biased algorithms and presented to people in addictive interfaces that overwhelmingly generate wealth to US elites, is the change that people want. " Trump Signs Previously Shelved AI Executive Order - summary of the EO. Widerstand gegen Kameras - German article about resistance against surveillance cameras; its history, methods and legal consequences. Person in the comments has an interesting tip: A brush, and acrylic paint mixed with sand. UN-Report zu KI-Umweltkosten - German article about how the UN had the chance of holding tech companies accountable in a new report, but instead only asks consumers to adjust their behavior. I am not opposed to also asking people to rethink their consumer decisions (otherwise, I would not resist using animal products, flying, getting a driver's license etc.; if there's no buyer, there's no product), but for the biggest impact, we need to focus on the source and hold companies to a high standard - or ban their business model or product entirely. The report was also seen as low quality by experts in the field(s). Appeals Centre Europe Transparency Report April 2025 - March 2026 - The Appeals Centre is an independent out-of-court dispute settlement body active due to the Digital Services Act; they've only been around for 18 months. If you are in the EU, you can use them to challenge social media platforms’ decisions on groups, pages accounts or other content which has been removed or kept up despite reporting (if it is about anything other than impersonation, hacked accounts, copyright or CSAM, but hopefully those too at some point). Most cases seem to be about account suspensions, nudity, fraud and scams. So far, they have processed more than 24,000 disputes, where 12,000+ of those fell within their scope. The report has some stats about their work, how many times they disagreed with the platform and overturned the result, and more. DSA User Support Guide also by the Appeals Centre; good breakdown of your rights under the Digital Services Act. The platforms are supposed to tell you that orgs like the Appeal Centre exist, but somehow still don't, and many people don't know their rights. Hold them accountable! Know your rights and make use of the newly established bodies. Under Article 20 of the DSA, users must be able to lodge complaints, free-of-charge, against decisions taken by the platforms within the last six months. Dark Patterns in AI Chatbots - self explanatory; basically about design and interaction/output choices that maximize usage and data collection, lie about the capabilities and emotional intimacy etc. I learned a new term: Privacy Zuckering! Also made me read this about Gemini encouraging a guy to kill others, steal a mannequin, and then kill himself. Arbeitspapier Identifizierbarkeit - German BayLfD summary and interpretation around identified and identifiable personal data in edge cases/gray areas, especially around pseudonymous data. What means to identify "count"? Not just your own! There's a difference between relative/subjective identification and absolute/objective identification. Sidenote: Love that they recommend RSS-Feeds or a Mastodon Account to keep up to date on legislation in this. From intent to action: the leaders' guide to building AI-powered workplace - paper sponsored by Kyocera and done by Economist Impact, based on a survey of 639 senior executives conducted in October and November 2025, with in-depth interviews with businesses and "thought leaders" in AI, digital transformation and workforce strategy. So... take it with a grain of salt, it is very corporate and very incentivized to be pro-AI in the workplace. Their key findings show that they want more investment, more adoption. But: Despite the "propaganda" (so to say), it exposes a lot of weaknesses everyone is already talking about in the workplace. To name one thing I scoffed at: Page 12, the fact that so many measure ROI of AI use in vague "employee productivity", which is probably just increased output or increased closed cases, without looking at the quality. Sad. 4% are not even measuring any ROI for it! Our Data After Us - paper by the CNIL about our digital remains. Covers questions like: Do you want the content to remain after your death? Who gets to have access and manage it? Should that person delete it, or should the platform automatically delete it? Should your remains be used to train an AI to impersonate you to help your loved ones? There seem to be age and gender differences to these answers. You Trust Your Chatbot With Everything - Should You? - paper by Theodore Christakis from AI-Regulation.com. The findings are as expected: Every major provider now trains on consumer chats by default, providers typically reserve safety and abuse-prevention uses and feedback actions to override the training opt-out, and they all reserve the right for humans to read the conversation. The author suggests a " Sealed Mode " where the default settings/options constrain reuse and human access, allows no training, has no advertising, little personalization, and cryptographic hardening. In my view, it could be a good first step, but I fear in practice, it would be bastardized, as meaningless and misleading as Incognito Mode in browsers has been. Ideally, the things of a Sealed Mode should be the default you can then opt out of one by one, and it can be legislated so. We have seen that hidden settings within different menus and specific modes you have to first know about and then turn on do not help the average user, since they are never actively prompted about them or told about them by the company. This stuff only aids a risk-transfer from controller to data subject. So do not offer a silly little compromise - make them default, and do not allow it to cost anything. Choosing between payment or privacy sucks. We should sometimes ask ourselves: If LLMs are just another tool, would I want Microsoft to always have access to and review my Excel sheets? Of course not! So why should we accept this here? At times, the author is too timid for me (" Yet the purpose of adopting this prism is not to export the GDPR as a universal template, nor to argue that the world should converge on European legal categories of individual control. " hey, why not? We don't have the Brussels Effect for nothing; privacy legislation worldwide has been shaped by the GDPR, one example being Brazil!). Favorite chapter was the second one (Ghost in the Machine), as it goes in on how incomplete and lacking the warning labels are, together with how contradicting they are when everything else encourages you to freely share anything. Least favorite are the parts where chatbots are asked to answer something; I am sorry, but I will never see these as genuine, truthful, verifiable answers. This is treating them as a conscious employee that an regurgitate internal policies, not a probability machine who can be nudged to give specific output. Gewalteskalation als System: Nihilistic Violent Extremism in Deutschland - German paper on NVE that's mostly done by children and teens, who connect online over misanthropic and nihilistic tendencies and then see extreme violence and vandalism as the only way forward. Not always far-right or incels, but often. The paper explicitly mentions the Com network, 764, MKY and NLM. Aside from Telegram, Discord is the biggest place for it. I was surprised how lax and wide the definition of violent extremism is (imo, that would make a significant portion of the population violent extremists), and I think the way the authors narrow it down a bit is a good attempt. 28.05.2026 – 26 O 869/26 aka the big one currently making the rounds about Google being responsible for the AI summary output. It will be interesting to see how that progresses and if it will be overruled. This one for noyb. Don't know if it counts as it is a web format, but I finished reading 17776 by Jon Bois.

0 views

Building confidence in geospatial data

How SkaldMaps generates a confidence score for data attributes that helps you gauge how accurate data is (or isn't).

0 views
Stratechery 2 months ago

The Nvidia AI PC, Project Solara, Microsoft AI

Listen to this post: Good morning, I don’t normally give away my interview subjects ahead of time, but I’m going to make an exception this week given the subject and the below Update. I am writing this in San Francisco where I interviewed Microsoft CEO Satya Nadella after his Build developer conference keynote ; normally I would want to publish that immediately so that you have the full context of my analysis. In this case, however, I came to the opinions below during the keynote, and before the interview, so for that reason (and a few logistical ones) I wanted to articulate them first (before you see my questions), and follow up with Nadella’s view on them (and a number of other topics) afterwards. So with that noted, on to the Update: From CNBC : Nvidia has emerged as the world’s most valuable company by dominating the market for artificial intelligence chips in the data center. Now the company is expanding its prowess to chips that will serve as the main processor for personal computers, entering an arena that’s long been ruled by Intel, Advanced Micro Devices, Qualcomm and Apple. During a keynote address at Taiwan’s Computex conference on Monday, Nvidia CEO Jensen Huang unveiled a new PC processor made alongside Microsoft. The RTX Spark superchip, which Huang also referred to as the N1X, debuts in the fall on a fresh line of Windows PCs from Microsoft, Dell, HP, ASUS, Lenovo and MSI. I’m actually starting in Taipei on Sunday, where Huang introduced the long-rumored Nvidia PC chip; from Tom’s Hardware : At full strength, this chip offers up to 20 Arm CPU cores, a Blackwell GPU with 6,144 CUDA cores, 128GB of LPDDR5X RAM, and up to 300 GB/s of memory bandwidth. That powerful CPU and GPU, connected over NVLink C2C, and the large memory pool give AI agents and 120-billion-parameter models plenty of power and space for long-running tasks with context lengths stretching to a million tokens, according to Nvidia. We don’t have any benchmarks yet, but the RTX Spark appears to be broadly similar to the DGX Spark; that’s a decent chip that excels at prefill, but is slower than an M5 Max at decode (thanks to lower memory bandwidth), and significantly slower at CPU tasks. Huang appeared during the keynote via live video to discuss the chip. Satya Nadella: Suddenly, this concept of unmetered intelligence right at the edge is so hot again. So maybe you want to talk a little bit about this: you have thought about this, talked about this, and now, of course, with RTX Spark really delivered, I think, what’s a breakthrough system for AI to be much more ubiquitous. But maybe, Jensen, you can just share a little bit your vision around where you see this going. Jensen Huang: Well, this all started about three years ago between a conversation between you and I. And we were talking about how we could build a new class of PCs that’s incredible for designers and creators. And it would be incredible for artificial intelligence. And it would be one of these systems that has the processing capability, but also the software stack that’s integrated into the world’s design packages and creator packages. And, of course, all the things that we’re doing with AI. And here we are, three years later, we built an incredible new chip. And this system is supported by all of this new software that you created for Windows. And we now have the ability to have essentially an autonomous agent running on the PC. This clip explains why I find this chip specifically, and AI PCs generally, pretty underwhelming. Three years ago we were still in the ChatGPT era of AI, and I was very excited about the possibility of local inference. Then came the reasoning era, blowing up KV cache (which increases the need for more memory) and emphasizing the importance of decode (to generate that many more tokens). Now we’re in the agentic era, where CPU performance is incredibly important. To that end, the ideal setup for a local agent is strong local CPU performance and calling out to the cloud for inference. The RTX Spark, however, spends tons of die space on GPU cores that are inferior to the cloud (because of memory size and bandwidth if nothing else) at the expense of CPU. It’s a suitable chip if you just want a chatbot circa 2023; it’s hard to see it being worth the price — or the software compromises that are the reality of Windows on ARM — in 2026. Jump ahead to the Build keynote, which I found very underwhelming to start. Nadella opened with a brief overview of the AI stack, then started talking about Windows, and I was honestly pretty surprised at the lack of vision and enthusiasm. That’s when it occurred to me: I think that Nadella agrees with me! Sure, some local inference is nice, but that’s not where the AI that matters is going to be located. Nadella, keep in mind, has no real loyalty to Windows; indeed, I credit him with The End of Windows . Specifically, Nadella didn’t end Windows as a product, but he ended its run as the organizing principle around which the entire company operated, focusing on software that ran everywhere and a cloud that ran everything. That leads to a surprising takeaway, and the most interesting part of the Build keynote: what if Microsoft is actually well positioned to get back into AI devices? From GeekWire : A team inside Microsoft has been quietly building a platform for devices that run AI agents instead of apps, based on Android instead of Windows, with two working hardware designs so far, and an initial set of big-name companies lined up to run pilots. The platform, dubbed “Project Solara,” is Microsoft’s bet that AI will open up entirely new scenarios for computing — using agents to avoid the constraints of traditional software, and off‑the‑shelf components to develop new devices quickly and inexpensively. Project Solara is, to be clear, vaporware at this point, although the company did show real devices and has signed up Qualcomm and MediaTek as chip partners. It is also extremely compelling. Here’s how Nadella introduced it: So far, we’ve talked about the edge and the cloud. The current form factors, right? I mean, when I saw that Jensen picture from the weekend where he had all the desktops, I felt like, man, I’m back in the 90s, right? Because it was so cool to see the lineup of all the machines that I loved and I grew up with back yet again with new functionality, right? It’s the same form factor, but unbelievable new functionality because of the onboard AI capability, right? So that’s sort of what we’ve seen with the laptop, the desktop, and of course with the cloud. But it also, you know, sets up that next question: if you have that capability, which is new function, and you can put it into existing form factors, can you even purpose-build new form factors for the new function? Can you build a new platform even for the agent era? And that is the motivation behind Project Solara, which we’re introducing today. First off, note the framing: the PC is old tech with agents; what about new tech uniquely enabled by agents? And note the classic Microsoft hook: could that new tech sit on top of a new platform? Corporate Vice President Steve Bathiche, the head of Microsoft’s Applied Sciences Group, explained the vision: Before I talk about those awesome new devices you just saw, let me start with the why. Back at Build 2023, I talked about the outside AI application structure, where AI moves from operating within the application frame to operating globally, working across multiple apps and services to connect, coordinate, and maintain context across entire workflows, devices, and time scales. What if there were an ecosystem of devices specifically designed for that new type of application structure, for those types of agents, for that transformational interaction technology? That is the impetus behind Project Solara. But with so many possible forms, which one do you pick? What is the next device? You see, the big aha for us is that it’s not about choosing one specific form factor. It is about creating a system that extends your agent across a constellation of devices. The next computer is not one device. It is all these devices working together as one system, with agents showing up closer to where and when you need them. There was one brief moment in the promotional video that preceded Bathiche’s appearance that made the concept click for me: The problem with wearable devices is the interaction model: they are only useful when you are interacting with them, when the human is in the loop, but being in the loop with a wearable is annoying and inefficient. What is being demonstrated here, however, is a brief interaction, and then an agent doing work in the background. In other words, the usefulness happens in the cloud without the human needing to be involved, because an agent is doing the work. That’s what I find compelling. On one hand, you can make the case that of course Microsoft would be interested in a device model that uses the cloud as a platform, given that Microsoft doesn’t control a mobile device like an iPhone. What occurs to me, however, is that even if Microsoft doesn’t succeed with Project Solara, this model — where the cloud is the hub and multiple devices are the spoke, instead of the phone being in the center — is clearly a better one for agents. Agents work best in the cloud, and across apps and devices; yes, the phone might be one of those devices, but when it comes to agents it shouldn’t be the hub. Again, this is vaporware, and very much in Microsoft’s interest, so take Project Solara with the appropriate grain of salt. It’s a vision of the future, however, that does make a lot of sense, particularly in an enterprise scenario where all of the context and compute is already in the cloud (and Project Solara is focused on enterprise, not consumer). It’s also something completely different from the past, and fits my thesis that, in the age of AI, thin is in . From GeekWire : Microsoft has based much of its AI business on models from OpenAI, before expanding more recently to Anthropic. On Tuesday, the company showed how it plans to rely less on both. At the Build developer conference, the Microsoft AI Superintelligence Team unveiled a family of seven models built from scratch. It’s part of an ongoing effort by the company to build credible in-house alternatives to models from partners and rivals with competing allegiances… The flagship of the seven newly announced MAI models is MAI-Thinking-1, a reasoning model that Microsoft says draws even with Anthropic’s Claude Sonnet 4.6 in blind human testing, and matches the more capable Claude Opus 4.6 on a widely used coding benchmark. [CEO of Microsoft AI Mustafa] Suleyman stressed that MAI-Thinking-1 was trained from the ground up with no distillation from other companies’ models, looking to appeal to enterprises that care about clean data lineage. These models seem pretty decent, all things considered, but what was interesting to me was the framing: Microsoft emphasized that enterprises could take these models and make them their own. Suleyman said: This is what owning the full stack end-to-end looks like. It’s the foundation of Microsoft Frontier Tuning, it lets you customize the MAI models using our full stack hill climbing machine right where you want it. And it means that the disciplined and very relentless engineering that has gone into building our models is now available to all of you on a platform that you can trust, working on your behalf to create custom agents that you will control. So the really big thing, of course, that’s happened in the last year is these RLEs, reinforcement learning environments, these unique training gyms for your AIs. They create company and task-specific agents adapted only to you, built on MAI models. So for example, within Microsoft, we use our RLEs combined with our MAI models to climb towards the best agentic use cases on Excel. Our MAI-tuned model is now on par with GPT 5.4 on public and private benchmarks, whilst at the same time being 10 times more efficient on cost, and many other early adopters are seeing similar results. When we’ve tuned our models on McKinsey’s tasks, MAI delivered the highest win rate, even outperforming GPT 5.5, and again delivering 10x greater efficiency on cost. So to us, this is the advantage of very carefully calibrated frontier tuning. And importantly, unlike with some of the other companies, with MAI, you don’t rent intelligence from a shared model that learns from everybody. Only you keep the benefits of your hard-earned workflows, know-how, knowledge, and your own institutional data. Only you get to control the resulting model. And so with us, the RLEs and the models that you build inside of them, they become your moat. I really think this is distinct. It marks a new era in AI that we’re all very, very excited about. This has shades of AWS’s Nova Forge offering , which lets enterprises add their data at a checkpoint in pre-training; it’s a little different in that it’s more focused on reinforcement learning, but those lines are getting blurred. The concept is that enterprises get to have their own model for their own data, without sharing it with the frontier labs that want to eat their lunch, and it’s a concept that is certainly appealing in theory; the real test will be to see if enterprises that choose this route aren’t penalized by not being on the cutting edge of functionality. Then again, helping cautious enterprises embrace the future on their terms, without necessarily having to win on pure performance, is exactly how Microsoft has long maintained its position. This Update will be available as a podcast later today. To receive it in your podcast player, visit Stratechery . The Stratechery Update is intended for a single recipient, but occasional forwarding is totally fine! If you would like to order multiple subscriptions for your team with a group discount (minimum 5), please contact me directly. Thanks for being a subscriber, and have a great day!

0 views
Aran Wilkinson 3 months ago

Introducing Headcode: A Unified API for UK Rail Data

Headcode is a unified, developer-friendly JSON API that takes the fragmented, legacy feeds of the UK rail network and turns them into clean, enriched real-time data.

0 views
Gabriel Weinberg 3 months ago

More data supports science funding literally pays for itself

Previously I put out a post explaining “ how science funding literally pays for itself ” that takes you through the math and some data that backs it up. Now two new data points further bolster this claim. First, the Congressional Budget Office (CBO), the nonpartisan federal agency that provides budget and economic information to Congress, published a report entitled “ Estimating the Economic Effects of Federal Investment in Research and Development . ” Usually the CBO only projects out 10 years per their mandate, but because the effects of science funding can take longer to fully manifest, they projected out 30 years. Thanks for reading Gabriel Weinberg! Subscribe for free to receive new posts and support my work. The relevant headline takeaway is highlighted below in their primary table (Table 1), showing that over this period the effects of a $30B increase in science funding for 10 years ($300B in total and about a 33% increase from today) would result in decreasing the overall deficit over 30 years (see green arrows). The decrease is about -2% on average if the “R&D funding increase [is] financed by reducing noninvestment spending” and about -1% on average if the “R&D funding increase [is] financed by borrowing.” This means that the increased science funding would grow the economy so much that the tax revenues received from this growth alone would outweigh the spending increase, leading to an overall decrease in the budget deficit. In other words, increasing science funding (at least by this amount) is a complete no-brainer, so let’s do it already! A few years ago the CBO did a similar report for infrastructure spending and compared the two in this report, finding the ROI effects of science funding to be about seven times greater than infrastructure spending. Again, so let’s do it already! The effect on the present value of GDP over the next 30 years (discounted using Treasury rates) that a dollar increase in deficit-financed R&D spending would have is about seven times larger than the effect that CBO, in its August 2021 report, estimated the same increase in infrastructure spending would have. Second, the Clark Center regularly polls a panel of economists , and recently they asked about this specific topic . The panel essentially universally agreed that historically U.S. science funding has paid for itself. In particular, 82% agreed “historical federal support for scientific research has paid for itself through a substantial positive effect on long-run U.S. productivity growth.” 0% disagreed, with the rest either not answering, or declaring either “no opinion” or “uncertain”. They also ask respondents about the confidence in their answer, and when weighted the results are even more striking with a whopping 97% in the agree category. Are you sold yet? Government science funding, the bulk of which goes to medical research, extends our lifespans and healthspans by inventing new medicines and other technologies that grow our economy so much it literally pays for itself. I get that this is not the most flashy policy area, but it is the most obviously good for our long-term future. Finally, and also new this year, the Pew Research Center put out a survey on Americans’ views of science and science funding , and among other things found broad bipartisan support for government science funding. 84% of U.S. adults say “government investments in scientific research aimed at advancing knowledge are usually worthwhile investments for society over time.” That breaks down by part as 76% of Republicans and 93% of Democrats (including independents who lean one way or the other). Thanks for reading! Subscribe for free to receive new posts or get the audio version .

0 views
ava's blog 3 months ago

beware of EU-washing

Among all this talk of European sovereignty and switching to European alternatives in a move to better privacy and less support of Big Tech, I wish for more emphasis on not just blindly copying US products and slapping an EU label on it. I see news like the Germany’s Federal Office for the Protection of the Constitution backing away from using Palantir and using a software solution from France instead. I’m supposed to feel happy reading this, and admittedly I did not yet dig into ArgonOS deeply - but all I can think of as a first reaction is “I don’t want an EU version of Palantir.” I don’t want ‘GDPR-compliant’ facial recognition and behavioral surveillance in our cities. I don’t want more privacy-friendly warfare (???). I don’t want more tech-enabled discrimination from next door. I don’t want supposedly European alternative that’s still based on AWS and Microslop. We need to be critical and take a stand against EU-washing, in which unethical business concepts or structures get painted in a more ethical light using the (increasingly less warranted) good reputation of the EU about human rights. We aren’t better for being from a different area, or just because it’s a different company name slapped on; it’s because we are supposed to have strong consumer protections and rights, resist the promise of easy money through unlimited data mining, and stand up against fascism. I don’t want us to compete with evil; I don’t want us to stoop to that level at all. Go hard on these copycats. Taking concepts from Fascism Land isn’t worthy of praise and they don’t deserve you as a customer or fan. Make them prove it first and ask them the hard questions. Boycott their shit if it is the same garbage, go to protests, write to representatives, be vocal online, support NGO’s that work against this. No one gets a pass for being European. I won’t lower my standards and values. Reply via email Published 24 May, 2026

0 views
ava's blog 3 months ago

computers, privacy and data protection conference 2026

I attended the Computers, Privacy and Data Protection Conference (CPDP) in Brussels for the first time. The conference has lots of different rooms mostly in the same building where multiple panels, workshops and other things are happening at the same time in specific slots, so you gotta choose what you participate in (was difficult at times!). Next to that, you have some fun rooms, some quiet working spaces and spaces to just hang out and talk. Based on the programme, the focus this year was definitely on age verification/youth 'protection', human AI relationships, consumer rights and marginalized groups. Lots of different groups and people present; people from the EU Commission and Parliament, AlgorithmWatch , Bits of Freedom , noyb and Max Schrems, IGLYO , EDRi , Equilabs , Equinox Initiative for Racial Justice , INTITEC , the EDPS and Wojciech Wiewiórowski, Privacy International , the International Committee of the Red Cross , the Office of the United Nations High Commissioner for Human Rights , the European Consumer Organization (BEUC), Future of Privacy Forum , AIRegulation.com , data protection authorities of different countries (CNIL, BFDI, etc.), ALTI , European Disability Forum , d.pia.lab , AI Now Institute , OECD , the IAPP , and all kinds of universities, plus companies like Mozilla, Mastodon, Signal, Wikimedia, Microslop, Uber, TikTok, Google and more. I was there for the opening remarks, then went on to visit: My takeaways/new things learned: Microsoft co-wrote parts of the EU's Energy Efficiency Directive , which allows data centers to keep their energy use confidential under the guise of business secrecy. The draft literally had paragraph's of Microsoft's proposal copied in unchanged. The Dutch government used racial/ethnic profiling via algorithms in the assessment of childcare benefit applications, which led to false allegations of fraud against thousands of families, particularly affecting those from ethnic minorities. I heard about this before, but learned more about it that day. To contest it all and defend democracy, we all need to train our AI literacy skills , support and have good tech journalism that questions and exposes it all (404media is, imo, a good example of what they meant), crafting and changing the social media narrative around AI and Big Tech, listening to affected people, demanding transparency via standards and audits etc. We cannot forget that officials know ; many of the effects we criticize are not accidents or side effects, they are the entire point. Like when tech predominately negatively targets marginalized communities, this is a bonus to people in power, and nothing to be fixed. Workers can resist by reminding their leaders of the liabilities and legal risks, strategic issues, money issues etc. that AI brings; demand specific definition of the needs that AI will fulfill at the workplace, instead of letting AI become the purpose instead of the tool. Age verification is racist and migrantphobic : Many people have issues with their ID, or have none, or are undocumented, and age verification in their country requires them to have contact with officials, police, etc. Age verification is transphobic : Relying on ID means many trans people are forced to reveal their deadname or are forced to come out, as it reveals they are trans if the ID is not or cannot be updated. The platforms are harmful, but we have so many ways and ideas against that that doesn't take away important spaces and support groups or bar entire groups of people. Age verification makes it possible for platforms to avoid working on their problems and becoming better, enables avoiding legislation and regulation, and enables control and surveillance by them; meanwhile, the truth is that you don't suddenly turn 16-18 and know how to handle porn, gore, harassment and all other negative parts of social media. The negative sides to social media that are named as the reason for age verification and banning of social media for specific age groups also affect adults negatively . We need to put more effort into education on how to handle these things. Yes, we can protect children's privacy by banning them off of platforms, but this also affects their other (digital and offline) rights, and privacy rights don't trump all . Children and teens should learn and be encouraged to control their own spaces and moderation via FOSS : Matrix, Mastodon, etc. where they can also seclude from adults and aren't reliant on Big Tech. Age verification and banning would take this away from them and also make it harder for FOSS projects. If children only ever enter the political discourse as victims, the only response can be rescue; that it why we have to make sure they enter as participants. Protection is not (just) space away from the risk, but confronting the systems that cause harm and eliminating them. 16-18% of US citizens report having engaged romantically with a bot, 45% of them said it made them feel more understood, 36% said it gave them stronger emotional support than their human partner. Problem: Current version of AI Act doesn't cover romantic and sexual use, no guidance for safeguards for emotionally responsive AI systems that protects around the risk of suicide, crimes, distress when service slows down or shuts down or model changes, discrimination as you get more if you pay etc.; drafts mention some of it now in Art. 50. With all the talk around becoming emotionally dependent on AI, nudging into harmful behaviors, etc. we cannot forget that you are also vulnerable on other services and in human romantic relationships, where the same routinely happens (weak argument, but to be fair, I also often forget this). We also cannot forget that it is not always a replacement - it often just supplements social life, and there are also surprisingly many people who just don't want or need romantic or sexual relations with a human ; they want bots specifically , and only bots. Disclosure agreements (meaning: labels everywhere that this is just a bot and not real) are most often useless, because people know and intentionally seek it out (exception for Insta/Snap DMs etc.) The latter about Human-AI intimacy was extra interesting because it had someone on the panel who directly works with people who use bots for romance and sex, and her experience has been mostly positive and that it helps her clients. Afterwards, I sadly was too overwhelmed, exhausted and in pain to continue and went back to the apartment to rest. Unfortunately, all the stress around the apartment and the generally more exhausting day triggered my digestive tract badly (Crohn's disease), but within the first few hours, all toilets in the venue were out of service due to an issue outside the venue or the organizer's control, and the alternative toilets were much further away. I didn't wanna have to deal with that with upset intestines. I missed the ' Designing Fairness ' Workshop, and the ' Consumer Rights at the age of acceleration' panel. Didn't meet anyone that day. Look at this ridiculous Gemini Photobooth they had that I saw no one use in the entire 3 days. This day, I managed to attend everything on my list, thankfully, as I felt a bit better. I attended: My takeaways/new things learned: The digital omnibus is mostly there to enable AI made in Europe to aid sovereignty and be competitive with US and China; AI here needs a framework to access data without much regulatory risk - that is what the EU Commission person said. Enforcing the law and and making it sharper is actually leveling the playing field and furthering innovation, because there is a massive power concentration of a handful companies that can do what they want, barely pay fines, have the fines suspended because of the US government bargaining with the EU, or who see them as a cost of doing business. Competition is impacted this way, as small companies are hit harder than the big ones. If the omnibus goes through with changing definitions of personal data etc., it will take years for case law, literature, standards etc. to catch up, it wastes money in companies who need to re-do everything to comply; so it doesn't simplify anything and makes praxis harder. You may set ChatGPT/Claude/Gemini etc. to not send feedback or training data in your settings, but when you react thumbs down/up to their request of whether the output was good or not, or choose between two different versions, the entire chat log until then gets sent for training and potential human review. So, these popup feedbacks override your settings . I need to read more papers by Theodore Christakis. Here is one of them. US and UK discovery and disclosure laws/principles go directly against EU data minimization principles; as long as data is relevant to a case it should be accessible, which is why in their cases, they can just have access to million's of people's data if necessary, and in a divorce case, they have the right to ask for AI chatlogs. There is no AI protection or privilege: If you use AI for legal stuff, you have no expectation of confidentiality like you would with a lawyer, so it is not safe from discovery. There is tension between tracking for harmful behavior/threats vs. data privacy rights ; what if someone threatens to kill themselves, kill others, etc.? Should company look for it, track it, report it, alert anyone, suspend the account, send help resources? Still unclear. There is also tension between people wanting the bonus features/ease of use coming from pesonalization and free services, while also not wanting to be tracked or charged. Advertisers see themselves as enablers of a good thing, as people want fitting ads, good algorithms, good suggestions, and free access; so if their business model is challenged or fails, people will have worse access and worse user experiences in their view. They also fear that if their business model is hindered, things will move into a more extreme, embedded, hard to avoid direction that you don't control or decide (Black Mirror ad type of stuff). I previously wrote about Consenter on the blog, and one panel had people from it there and showing screenshots; changed my mind on it a lot and made me understand the new features and goal better, I will probably write an update on it some time. We have different other options all covering something different about tracking, cookies, consent, or going about things differently, old and new: ADPC, GPC, ConStand, Global Privacy Control, DoNotTrack etc.; important for new stuff is granular consent, sent to the website, user given explanations etc. Uninformed decisions and bad practices lead to unfair competition ; bad actors erode trust level overall, so users resignate, experience fatigue and say yes in the same rates between "good" and "bad" services. Will read soon: Our data after us by the CNIL , and future release: Model rules on succession and access to digital remains by Eigenmann und Harbinja Digital remains can be split into assets (copyright, crypto, business tools, money), personal (messages, photos, identities, AI replicas), and third party data. GDPR only addresses living people; dead people's digital remains are subject to member state laws. There might be a need for something harmonized and European, though. For good digital hygiene , we should remember death and make it as easy as possible or sensible for the people we leave behind to get the access they need to manage our stuff how we want them to. Leave instructions, set emergency/legacy access when available (Google, Facebook, Instagram and Apple have it), include digital assets in your will, decide how your data is allowed to be used after death, especially around AI replicas. Hospice, nurses, families etc. should learn to ask affected parties about these things. Thanks to the focus on agentic AI, there is massive need for inference compute, which is super expensive. Almost all of it is in the control of, or can only be afforded by, the hyperscalers. At the same time, anything that seeks to enable or disable things for AI agents on the web can also affect accessibility programs like screen readers. It is in the best interest of the Big Tech companies to keep things individual, because it distracts from the collective issues and changes they'd have to do; it is easier to blame the person for agreeing to tracking than make sweeping changes to how much can be tracked. Individual consent doesn't consider the fact that data doesn't just affect you, but reveals things about your family, friends, partners, coworkers and more, as data is deeply interconnected. If your friend agrees to share his data and it also includes you, that is your data, still going to the service you'd have disagreed to. We as users have no collective bargaining tools yet; even big worker unions aren't negotiating with Microsoft about the terms of their employer using Microsoft Teams, when they actually should. We should also build up data unions made from users who bargain with the platforms. Strikes could look like boycotting the service, blocking trackers, scrambling data, massive amounts of access requests etc. Look into something called a Worker Data Trust ; this was used to prove Uber's predatory dynamic pricing (Worker's Info Exchange). Lots of workers made access requests, the data was combined and analyzed by researchers. After a failed attempt to meet up during lunch, I managed to meet up with another Country Reporter from noyb for a little while until the next panel happened, and sadly we didn't go to the same one. At this point, I was miffed about lunch at the conference. They made a big deal at registration about how the event will be mostly vegan and vegetarian to offset the climate impact of everyone traveling there, and they asked you to select your preference. I chose vegan. But for the entire three days, the food wasn't clearly labeled, some food was mislabeled as vegan when it wasn't, and there was way too little of it and wasn't restocked. It was more like "vegetarian snacks for birds". Vegan people had no warm food option at all, just sandwiches or wraps all three days that would have been enough for maybe 10 people. I mostly starved and I accidentally ate real cheese one time too because the food situation was so confusing. Here was one of the buffet menu cards, which were a bit to the side removed from the food, partially hidden by other stuff, and incorrect (anything with lactose is not vegan). I have no idea how, on a sea of silver platters with lots of bread, I am supposed to be able to differentiate the vegan gluten free bread option and the vegetarian gluten free bread that has scarmoza (italian cheese). It was a roundtable buffet, so everyone was waiting on you to hurry and grabbing stuff; I can't just grab bread and lift off the top to see the ingredients and then put it back, man. At least group the vegan stuff together or put labels directly in front of each thing. Also, while I am not reliant on gluten-free food, I think the people sensitive to it or having celiac disease don't appreciate that either. I skipped the Cocktail parties and big CPDP party, because it's not really feeling fun when you don't drink alcohol, have trouble just going up to people with your mask and hoping they hear you, and have no one to meet or go with. Last day was rather empty in the programme, so I arrived later and left earlier. I attended: My takeaways/new things learned: The AI warfare one was a bit of a letdown, because they all just accepted war as a right, an inevitable thing that has to happen. There was not even a nuance of fighting war itself, or banning AI weapons, etc; it focused more on the dual nature of the data , in which through surveillance, tracking, etc. not only can military use it to target people, NGO's and others can use it to warn, evacuate, render humanitarian aid etc. and document realities on the battlefield. There was also no possibility for the idea that we could enter an age where drones fight drones automatically and no one needs to get hurt or be traumatized or get to kill people like a game, and that is only because everyone is so attached to the idea that war has to have human casualties. It's hard to legislate and restrict because the data is taken from a whole ecosystem : Telecommunications, cloud services, civilian infrastructure, social media etc. and most of the data is collected during times of peace. Warfare is often explained with national security as a reason, which then again is a legitimate interest or fulfills other opening clauses in data protection and privacy laws. It is a problem that the richest men in the world, close to the US admin, lead the biggest companies worldwide, almost all in the US, and control almost all of AI and AI warfare. Project Maven from 2017 was continuously developed on and is now the Maven Smart System , which was used in Venezuela and Iran recently. Our Art. 15 GDPR right of access as it is right now is making up for Germany and Austria's lack of discovery and disclosure rights respectively. Controllers can usually drag stuff out, cite trade secrets and rights of others to evade data access, but the data subject barely has any power. Not having to justify the access request and it not having to be limited to data protection rights is good in this regard and needs to be kept up. Otherwise, also too much confusion and court cases whether a request was abusive or not if now, any request for a court case instead of privacy rights is deemed possibly abusive. We don't only need to focus on reidentification in general, but about the ability to single people's data out; you might not be able to identify them, but you can build a profile anyway. Learned about the term digital twin , or in terms of user data, a data twin that can be used for similation and is similar enough. AI-act-standards.com exists. Many don't know that the AI Act isn't a GDPR for AI, but serves more as market classification, as it sorts AI into different boxes who have to fulfill different requirements. The details of these requirements are/will be set with CEN/ISO standards and frameworks . You can see the progress of development on these standards on that website, and what they cover and how they interact. Hovering over the elements gives additional info. This is done by the JTC21 , and you can also get involved by registering with your national standardization body (in Germany, this is DIN) or when they do public consultations. Disabled people experience both extremes of AI - better accessibility options, often more reliant on AI, so also more subject to surveillance and having their privacy rights violated, while bad governments can use the data to harm disabled people, all under the guise of research. Marginalized groups are often the first trial group in anything, while not being stakeholders in the tech, or even invited to the table. See: AI used in immigration etc. and with deregulation and AI everywhere, we see a loss of reasonable suspicion thresholds in law enforcement and other groups. Learned about adversarial auditing . The previous two days, I did the whole fancy dress pants and blazer thing (one black blazer, one dark red/purple blazer), but for the last day and the drive home, I wore my Bearblog shirt and wide orange jeans: Someone from noyb staff thankfully recognized me and approached me, so we talked for a bit until he had to leave for another lunch meeting. That concludes the human contact I had. And then I left to drive home with my wife. She will hopefully soon write a guest post on my blog about how she navigates a new city in another country without mobile data/a smartphone (she has a tablet with WiFi only), because while I was at the conference, she explored the city on her own. It's kind of difficult to show up to these conferences as someone who isn't sent there for work, who doesn't have coworkers or ex-coworkers also attending, and who doesn't have much or any industry contacts yet. Most people there know each other from work or previous/other conferences, and I don't. These events are primarily for networking, keeping in touch, and talking about what you have seen and learned though. I couldn't discuss anything with anybody present, and it made me feel really lonely and silly. Just going up to people and striking up a conversation is not my strong suit, and it's something I am working on and has already gotten better, but the mask I am usually wearing in these big crowds and gatherings because I am on immunosuppressive medication is actively keeping me isolated. I know people have trouble understanding me, can't see me smiling at them, and think I am sick, so that keeps both sides hesitant. Unfortunately, if I attend next year, I will have to leave away the mask and maybe try out these protective sprays for nose and throat that are supposed to reduce viral load. It seems like you can only 'afford' to wear a mask if you are already in a group of people. Weeks before the event, I asked some people if they would attend, they said they will and we had a group chat of 10 to coordinate meetups. But during the entire conference, I was the only one trying to make something happen - saying where I am/where I will be, identifiers you could spot me with (as we never met before and you can't see name tags well on the lanyard), meeting points etc. and the two people mentioned were the only ones who took me up on it. The others just ghosted me/ignored my messages. That saddened me a lot during the conference. And unfortunately, these types of events are always really exhausting to me beyond the normal amount everyone experiences, because of things that trigger my conditions, my lower energy, my needs to lie down sometimes, sensory issues, food restrictions etc. so I really have to weigh if it's worth it to me. I'm not sure it is, without the social aspect. Many of the panels I chose had an issue of being not well organized. Instead of short speaker times, precise audience questions, interactions, dialogue, disagreements, different sides, answering the panel's topic and offering solutions etc., it often resulted in every speaker having a 10 minute monologue saying their peace, the other speakers not reacting or intervening because it's too much, everyone more or less saying the same thing or zoning out, and then having too little time to really give much attention to audience questions. Some gathered audience questions to answer them in batches and predictably, that resulted in nuance being lost and almost nothing being precisely answered. From many panels, I walked away with less learned than I wanted to, and just being reaffirmed in what everyone knew already. There were almost no further or new resources, or real takeaways of what the next steps should be and how we can tackle or solve an issue. They say " there should be more transparency " but not how we ask for it, how we legislate it, how it should happen. It's often just a vague " Someone should do more of something, and fast. " It was easy for people from the EU Commission to dodge mine and others' questions about the omnibus bullshit with no convincing answer. (: It disillusioned me a bit about my own goal to be speaking at a panel one day, because so often it felt like it was just there to platform someone to give them a chance to ramble and that's it, or just so that they can put this on their CV. Looking into the panelists, so many of them are genuinely great, very accomplished and admirable people with a lot of expertise, but the way things were set up, it couldn't shine through. You would have been better off talking to them directly. As a final bonus for reading this far, help me delete this (fortune) cookie. Reply via email Published 23 May, 2026 Contesting AI & Defending Democracy ; Possibilities for European AI Futures ( x ) Youth protection through inclusion and empowerment : a rebuttal of the exclusion-based narrative ( x ) Intimacy by Design: Governing Human AI relationships ( x ) Microsoft co-wrote parts of the EU's Energy Efficiency Directive , which allows data centers to keep their energy use confidential under the guise of business secrecy. The draft literally had paragraph's of Microsoft's proposal copied in unchanged. The Dutch government used racial/ethnic profiling via algorithms in the assessment of childcare benefit applications, which led to false allegations of fraud against thousands of families, particularly affecting those from ethnic minorities. I heard about this before, but learned more about it that day. To contest it all and defend democracy, we all need to train our AI literacy skills , support and have good tech journalism that questions and exposes it all (404media is, imo, a good example of what they meant), crafting and changing the social media narrative around AI and Big Tech, listening to affected people, demanding transparency via standards and audits etc. We cannot forget that officials know ; many of the effects we criticize are not accidents or side effects, they are the entire point. Like when tech predominately negatively targets marginalized communities, this is a bonus to people in power, and nothing to be fixed. Workers can resist by reminding their leaders of the liabilities and legal risks, strategic issues, money issues etc. that AI brings; demand specific definition of the needs that AI will fulfill at the workplace, instead of letting AI become the purpose instead of the tool. Age verification is racist and migrantphobic : Many people have issues with their ID, or have none, or are undocumented, and age verification in their country requires them to have contact with officials, police, etc. Age verification is transphobic : Relying on ID means many trans people are forced to reveal their deadname or are forced to come out, as it reveals they are trans if the ID is not or cannot be updated. The platforms are harmful, but we have so many ways and ideas against that that doesn't take away important spaces and support groups or bar entire groups of people. Age verification makes it possible for platforms to avoid working on their problems and becoming better, enables avoiding legislation and regulation, and enables control and surveillance by them; meanwhile, the truth is that you don't suddenly turn 16-18 and know how to handle porn, gore, harassment and all other negative parts of social media. The negative sides to social media that are named as the reason for age verification and banning of social media for specific age groups also affect adults negatively . We need to put more effort into education on how to handle these things. Yes, we can protect children's privacy by banning them off of platforms, but this also affects their other (digital and offline) rights, and privacy rights don't trump all . Children and teens should learn and be encouraged to control their own spaces and moderation via FOSS : Matrix, Mastodon, etc. where they can also seclude from adults and aren't reliant on Big Tech. Age verification and banning would take this away from them and also make it harder for FOSS projects. If children only ever enter the political discourse as victims, the only response can be rescue; that it why we have to make sure they enter as participants. Protection is not (just) space away from the risk, but confronting the systems that cause harm and eliminating them. 16-18% of US citizens report having engaged romantically with a bot, 45% of them said it made them feel more understood, 36% said it gave them stronger emotional support than their human partner. Problem: Current version of AI Act doesn't cover romantic and sexual use, no guidance for safeguards for emotionally responsive AI systems that protects around the risk of suicide, crimes, distress when service slows down or shuts down or model changes, discrimination as you get more if you pay etc.; drafts mention some of it now in Art. 50. With all the talk around becoming emotionally dependent on AI, nudging into harmful behaviors, etc. we cannot forget that you are also vulnerable on other services and in human romantic relationships, where the same routinely happens (weak argument, but to be fair, I also often forget this). We also cannot forget that it is not always a replacement - it often just supplements social life, and there are also surprisingly many people who just don't want or need romantic or sexual relations with a human ; they want bots specifically , and only bots. Disclosure agreements (meaning: labels everywhere that this is just a bot and not real) are most often useless, because people know and intentionally seek it out (exception for Insta/Snap DMs etc.) Simplification for Whom? Unpacking the Consumer Impact of the Digital Omnibus ( x ) My Chatbot, My Confidant: Protecting User Privacy in Generative AI Conversations ( x ) Informed consent: The breakthrough in Art. 88b GDPR / Digital Omnibus and current initiatives in the field of PIMS and technical standardisation ( x ) Digital Legacy Beyond GDPR: Succession, Data Protection, Access Rights, and Platform Power ( x ) The Agentic Assistant: What does Big Tech’s goal of creating a universal digital intermediary mean for society? ( x ) Designing Collective Technology Governance ( x ) The digital omnibus is mostly there to enable AI made in Europe to aid sovereignty and be competitive with US and China; AI here needs a framework to access data without much regulatory risk - that is what the EU Commission person said. Enforcing the law and and making it sharper is actually leveling the playing field and furthering innovation, because there is a massive power concentration of a handful companies that can do what they want, barely pay fines, have the fines suspended because of the US government bargaining with the EU, or who see them as a cost of doing business. Competition is impacted this way, as small companies are hit harder than the big ones. If the omnibus goes through with changing definitions of personal data etc., it will take years for case law, literature, standards etc. to catch up, it wastes money in companies who need to re-do everything to comply; so it doesn't simplify anything and makes praxis harder. You may set ChatGPT/Claude/Gemini etc. to not send feedback or training data in your settings, but when you react thumbs down/up to their request of whether the output was good or not, or choose between two different versions, the entire chat log until then gets sent for training and potential human review. So, these popup feedbacks override your settings . I need to read more papers by Theodore Christakis. Here is one of them. US and UK discovery and disclosure laws/principles go directly against EU data minimization principles; as long as data is relevant to a case it should be accessible, which is why in their cases, they can just have access to million's of people's data if necessary, and in a divorce case, they have the right to ask for AI chatlogs. There is no AI protection or privilege: If you use AI for legal stuff, you have no expectation of confidentiality like you would with a lawyer, so it is not safe from discovery. There is tension between tracking for harmful behavior/threats vs. data privacy rights ; what if someone threatens to kill themselves, kill others, etc.? Should company look for it, track it, report it, alert anyone, suspend the account, send help resources? Still unclear. There is also tension between people wanting the bonus features/ease of use coming from pesonalization and free services, while also not wanting to be tracked or charged. Advertisers see themselves as enablers of a good thing, as people want fitting ads, good algorithms, good suggestions, and free access; so if their business model is challenged or fails, people will have worse access and worse user experiences in their view. They also fear that if their business model is hindered, things will move into a more extreme, embedded, hard to avoid direction that you don't control or decide (Black Mirror ad type of stuff). I previously wrote about Consenter on the blog, and one panel had people from it there and showing screenshots; changed my mind on it a lot and made me understand the new features and goal better, I will probably write an update on it some time. We have different other options all covering something different about tracking, cookies, consent, or going about things differently, old and new: ADPC, GPC, ConStand, Global Privacy Control, DoNotTrack etc.; important for new stuff is granular consent, sent to the website, user given explanations etc. Uninformed decisions and bad practices lead to unfair competition ; bad actors erode trust level overall, so users resignate, experience fatigue and say yes in the same rates between "good" and "bad" services. Will read soon: Our data after us by the CNIL , and future release: Model rules on succession and access to digital remains by Eigenmann und Harbinja Digital remains can be split into assets (copyright, crypto, business tools, money), personal (messages, photos, identities, AI replicas), and third party data. GDPR only addresses living people; dead people's digital remains are subject to member state laws. There might be a need for something harmonized and European, though. For good digital hygiene , we should remember death and make it as easy as possible or sensible for the people we leave behind to get the access they need to manage our stuff how we want them to. Leave instructions, set emergency/legacy access when available (Google, Facebook, Instagram and Apple have it), include digital assets in your will, decide how your data is allowed to be used after death, especially around AI replicas. Hospice, nurses, families etc. should learn to ask affected parties about these things. Thanks to the focus on agentic AI, there is massive need for inference compute, which is super expensive. Almost all of it is in the control of, or can only be afforded by, the hyperscalers. At the same time, anything that seeks to enable or disable things for AI agents on the web can also affect accessibility programs like screen readers. It is in the best interest of the Big Tech companies to keep things individual, because it distracts from the collective issues and changes they'd have to do; it is easier to blame the person for agreeing to tracking than make sweeping changes to how much can be tracked. Individual consent doesn't consider the fact that data doesn't just affect you, but reveals things about your family, friends, partners, coworkers and more, as data is deeply interconnected. If your friend agrees to share his data and it also includes you, that is your data, still going to the service you'd have disagreed to. We as users have no collective bargaining tools yet; even big worker unions aren't negotiating with Microsoft about the terms of their employer using Microsoft Teams, when they actually should. We should also build up data unions made from users who bargain with the platforms. Strikes could look like boycotting the service, blocking trackers, scrambling data, massive amounts of access requests etc. Look into something called a Worker Data Trust ; this was used to prove Uber's predatory dynamic pricing (Worker's Info Exchange). Lots of workers made access requests, the data was combined and analyzed by researchers. Data-driven warfare : AI, civilian risks, and corporate responsibility ( x ) Digital Omnibus meets the Charter of Fundamental Rights ( x ) Toward a Standard for Fair AI-driven Recruitment ( x ) Data protection law as a shield, not a weapon: empowering historically marginalized communities in the EU in times of de-regulation ( x ) -> this choice was especially rough, because I was also very interested in ' The U.S. Deregulatory Effect ' happening elsewhere at the same time The AI warfare one was a bit of a letdown, because they all just accepted war as a right, an inevitable thing that has to happen. There was not even a nuance of fighting war itself, or banning AI weapons, etc; it focused more on the dual nature of the data , in which through surveillance, tracking, etc. not only can military use it to target people, NGO's and others can use it to warn, evacuate, render humanitarian aid etc. and document realities on the battlefield. There was also no possibility for the idea that we could enter an age where drones fight drones automatically and no one needs to get hurt or be traumatized or get to kill people like a game, and that is only because everyone is so attached to the idea that war has to have human casualties. It's hard to legislate and restrict because the data is taken from a whole ecosystem : Telecommunications, cloud services, civilian infrastructure, social media etc. and most of the data is collected during times of peace. Warfare is often explained with national security as a reason, which then again is a legitimate interest or fulfills other opening clauses in data protection and privacy laws. It is a problem that the richest men in the world, close to the US admin, lead the biggest companies worldwide, almost all in the US, and control almost all of AI and AI warfare. Project Maven from 2017 was continuously developed on and is now the Maven Smart System , which was used in Venezuela and Iran recently. Our Art. 15 GDPR right of access as it is right now is making up for Germany and Austria's lack of discovery and disclosure rights respectively. Controllers can usually drag stuff out, cite trade secrets and rights of others to evade data access, but the data subject barely has any power. Not having to justify the access request and it not having to be limited to data protection rights is good in this regard and needs to be kept up. Otherwise, also too much confusion and court cases whether a request was abusive or not if now, any request for a court case instead of privacy rights is deemed possibly abusive. We don't only need to focus on reidentification in general, but about the ability to single people's data out; you might not be able to identify them, but you can build a profile anyway. Learned about the term digital twin , or in terms of user data, a data twin that can be used for similation and is similar enough. AI-act-standards.com exists. Many don't know that the AI Act isn't a GDPR for AI, but serves more as market classification, as it sorts AI into different boxes who have to fulfill different requirements. The details of these requirements are/will be set with CEN/ISO standards and frameworks . You can see the progress of development on these standards on that website, and what they cover and how they interact. Hovering over the elements gives additional info. This is done by the JTC21 , and you can also get involved by registering with your national standardization body (in Germany, this is DIN) or when they do public consultations. Disabled people experience both extremes of AI - better accessibility options, often more reliant on AI, so also more subject to surveillance and having their privacy rights violated, while bad governments can use the data to harm disabled people, all under the guise of research. Marginalized groups are often the first trial group in anything, while not being stakeholders in the tech, or even invited to the table. See: AI used in immigration etc. and with deregulation and AI everywhere, we see a loss of reasonable suspicion thresholds in law enforcement and other groups. Learned about adversarial auditing .

0 views