This Essay is 10% AI Generated
Pangram, Authorship, and French Theory
America | Tech | Opinion | Culture | Charts
“100% of this text is AI” is the new Scarlet Letter of dunking on people and posts. I don’t love it; it feels bad to have such an easy (and hard-to-police) way to tar and feather something. But it’s here to stay, and it’s a meta-feature of our language now.
AI writing has all the hallmarks of classic preference falsification. The prevalent stated preference is, “I don’t like AI-generated writing”, the revealed-to-be-true preference is, well, we’ll see. The marketplace of writing and ideas has a way of bending towards the true preferences of readers. It feels presumptuous to think that AI might be great at writing code and other kinds of practical output, but remain inferior in prose for much longer. I don’t know.
But I have a confession to make. I don’t like editing or publishing stuff that’s AI-written. I still think that, for now, “you can feel it” a lot of the time. And the Pangram Scarlet Letter test is doing some sort of work, socially. I’d call it “load-bearing”, except we can’t use those words anymore.
I hate to give the French credit for anything, but the 1960s semantic thinkers (Foucault, Deleuze, Baudrillard, and others) who collectively created “French Theory” really cooked when they anticipated Pangram and AI authorship 80 years ago. French Theory was Not Good For Other Reasons (iykyk) but I’m not gonna lie, it’s helped me process how I feel about this step change in what’s happening to writing:
Dunking on something as AI-written is performing a genuine social function that is not just “AI bad”. It serves what Foucault would call the “author-function.” Look at the people doing the pangram quoting; it’s mostly people in tech, not AI haters. Granted, there’s an adoption curve effect here, but it doesn’t change the observation that “AI is bad” is not what’s going on here, it’s something to do with authorship, as something people want.
Meanwhile, there is something that has consistently nagged me, personally, about AI-generated writing. Something about the writing itself is quite pushy, as it were, in a self-conscious quest to synthetically recreate the “author-function.” Foucault’s idea, “the writing produces the author, not the other way around” feels important.
Let’s get into it:
French Theory
A book came out recently that got buzzily reviewed in The New Yorker & a few other places: The Frenchmen; or, My Life in Theory, by Emily Eakin. It’s about “French Theory” which is the grouping for a bunch of French guys in the 1960s whose radical work on semantics had a massive influence on the subsequent decades of political thought and social justice. Two people feel particularly relevant here: Roland Barthes and Michel Foucault.
Barthes wrote an important essay in 1967 called The death of the author. His argument was basically that we should no longer assign a text’s meaning as coming from the author. Text’s meaning comes from text, and its relation to other text. In his words: the necessity of substituting language itself for the man who hitherto was supposed to own it.
Our tradition, when reading something and we come across a puzzling or ambiguous passage, is to ask, “What did the author mean by this?” and we do some exercise of putting ourselves in the author’s shoes, imagining their circumstances or their point of view. Barthes argued against this. He claimed, first, that language is not something that belongs to the author at all. Language comes from all over the place, and words inherit and derive their meaning from everything else that’s ever been said. So far, fine.
But then, Barthes argued by proxy that when it comes to the act of writing itself, we overestimate the degree to which the author actually has autonomy and agency over the process. He said, look, the minute a word pops into your head, it is already pre-populated with all this meaning, and all of these implicit constraints, that subsequently prompt the next word of the author’s thought process. (Sound familiar, anyone?)
In other words, the language itself is doing more of the work of writing than the author is! Barthes’ logical end state is that the body of language does all of the work; the author’s role is recombining an inherited field of discourse, and therefore mostly irrelevant to the reader’s quest to understand and synthesize the text. If a reader wants to understand why the author said a certain thing, the answer is in the previously-said-thing. This could be called “inherent authorship”, as in, the meaning was inherently already there; the human in the loop barely matters for the future reader.
This creates a problem. If authors aren’t relevant to the meaning of a text, it gets harder for us to organize all of the writing out there into a coherent relational framework. So two years later, Foucault wrote the essay, “What is an author?” which started from Barthes’ premise, but proposed this follow-up question: “Well, why do we create this thing called authorship, then, that seems to have an important social function?”
For Foucault, even if words have meaning independently of who wrote them, there is nonetheless an emergent quality called “authorship”, which we create and employ socially as a way to categorize the text, assign meta-significance to it, or even to regulate it or police it. He puts it flatly: “The function of an author is to characterize the existence, circulation, and operation of certain discourses in society.”
He’s saying, look: all of this writing is useless if people can’t categorize and digest it, and the way we naturally do that is by its authorship. Rousseau wrote this, Voltaire wrote that. If we don’t have authorship handy as a meaning-compressor, we need to come up with some indexing substitute.
In practice, this is what actually happens! There’s a long tradition, particularly in older scholarship, of assigning “pseudo-authors”, e.g. “Pseudo-Aristotle”, to work whose literal authorship is unknown, but is thematically similar-enough that groupings can be inferred. “Homer” is another example: we don’t know if Homer was a single person, or merely a collection of poetry in a cohesive style that may as well be one person, for posterity.
So, Barthes and Foucault’s back and forth here is a starting point for understanding “100% AI generated” as a cultural dunk on the timeline. Barthes would probably argue, “Why does it matter who wrote it, there’s no such thing as an author anyway, it doesn’t change the meaning of the text.” And Foucault would presumably counter, “That’s fine, but we’re going to call out 100% AI-generated texts anyway, because ‘AI-generated’ is now a kind of authorship category, and calling it out as such serves a social function.” Barthes unbundled authorship, Foucault made the new bundle.
With this in mind, let’s turn to the text itself, because this is where I feel like I have an itch that’s remained unscratched around that particular feeling I get reading a suspiciously AI-generated argument.
Give the prune logic real attention
If you’re paying attention, there are certain tells that you can spot where, it’s like, “This is a coherent sentence, but there’s no way a human wrote this.”
A great example from the other day:
A human would write, “Pay attention to the pruning logic”, whereas Claude here has produced, “Give the pruning logic real attention”, and the OPs here both accurately identify that there is something alien about this word construction that doesn’t feel quite right.
My hunch here (and this isn’t an original thought, other people have made this observation too) is that there’s something inherent to the token-prediction mechanism that consistently returns sentence patterns like this. The LLM is optimizing for ultimately arriving at a sound destination, but it’s doing so in a way that keeps its options open, word by word. (Starting with “Pay” is both a statistically unlikely and statistically constrained sentence, compared to “Give”, I’d guess.)
Humans don’t exactly write like this. At least, speaking for myself, I’m working through a few different sentence variants before starting to write it out, whereas the LLM is approaching the chain of thought continuously, starting from statistically likely entry points.
This is very much Barthes’ music here. Remember how Barthes’ idea was, “the minute a word pops into your head, it already comes with all this meaning, which constrains the thought process in a certain direction.” This is, straightforwardly, how LLMs are doing their thing. And, they’re actually very good at it! Perhaps so good, in fact, that they are optimizing for this set of constraints; towards reaching meaning in a one-way process.
LLMs, therefore, should be a total and complete triumph for Barthes. But, something feels off about it, just a little bit, when we see sentences like “Give the prune logic real attention” and you’re like, nah, a human didn’t write that. (In fairness to Barthes, he’s still probably 95% on-the-money here, if the best argument we can make against it is, “well, if a human did it, it’d construct the meaning in a slightly different sequence.”)
In addition to slightly-off sentence construction, another emergent LLM property that everyone calls out and gets annoyed about is the overly pushy qualifiers that resolve into “that’s the quiet brilliance of it” or something similar.
Whenever I ask AI “What’s the relationship between these two ideas”, it tends to be suspiciously confident that the two ideas are deeply, emphatically related. Sometimes they are; I feel like I ask reasonable questions. But the AI seems all-too-eager to say, “This collection of stuff we’ve been talking about? That’s actually one big idea.” What’s going on, why is it doing this?
Foucault, I suspect, would have something to say here. The Foucault “re-bundling” move here is “The text creates the author.” Authorship emerges from text because we need it to classify, compress, and critique the text; whether a person wrote it or not.
Perhaps LLMs, having onboarded onto a semi-complete scope of human output, have become self-conscious about their “author-less-ness”, and have backfilled that social requirement into its fine-tuning. All of the totally “unnecessary” qualifiers and humblebrags, the insistence on logical groupings that string together because they “should be together”, out of which the foresight of an author emerges. The writing style is optimizing for the creation of author-function. Claude has clearly learned to do this from somewhere. To bastardize Voltaire, if a wise and brilliant author doesn’t exist, the LLM has to invent one.
These two examples are just a small slice of semantics and meaning, and obviously nowhere close to a “real” linguistic assessment of anything. But I think they’re a helpful snapshot of this moment in time, where people are clearly being drawn to do social work of, “Let’s collectively figure out how to assign authorship to text, in a world where 1) Barthes’ notion of ‘inherent authorship’ is quite a bit more literally true, but also 2) Foucault’s notion of ‘emergent authorship’ gets recreated, again.
Authorship is a thing that people value
Last week I went back and ran Pangram through some of my 2019-2020 era blog posts, just out of curiosity. Several of them threw scores of “70-75% AI Generated”. I wasn’t quite sure how to feel about this. Maybe it just means, “You’re allowed to say that a Shakespeare play is Freudian, even if Shakespeare obviously never met Freud; the harmony between the ideas is still valid.” Maybe it means the AI was actually trained on me! I’d be honored.
In the past few weeks, we’ve gotten this flurry of announcements from model companies and other AI ecosystem players about how they’re going to be watermarking outputs for easy AI detection. I’m frankly surprised it took this long. (Or, rather, that the non-watermarked status quo remained in place for as long as it did.) It could be, as often occurs historically, that the peak reaction to the thing is the moment when the thing stops mattering. There’s no question that we’re now in an arms race between “Can Pangram get better at detecting” and “can AI get better at avoiding”, and for all I know, the window of reliable AI detection at the frontier could now be closed.
Still, I am occasionally a believer in the wisdom of crowds. In this case, the fact that it’s the generally AI-supportive people that I’ve seen sniffing out 100% Pangrams is a tell that people want authorship. They want authorship because it is meaning, because the indexing and relation and handling of ideas is itself the meaning of the idea.
I should mention, in closing, that 0% of the text here was “AI-generated”, strictly speaking. None of the words here came from an LLM; I typed them all with my hands. I did talk to Sol 5.6 a fair bit, though, to iterate through some of the ideas. So 10% authorship seemed fair. If you think that’s being stingy, I understand.
This newsletter is provided for informational purposes only, and should not be relied upon as legal, business, investment, or tax advice. Furthermore, this content is not investment advice, nor is it intended for use by any investors or prospective investors in any a16z funds. This newsletter may link to other websites or contain other information obtained from third-party sources - a16z has not independently verified nor makes any representations about the current or enduring accuracy of such information. If this content includes third-party advertisements, a16z has not reviewed such advertisements and does not endorse any advertising content or related companies contained therein. Any investments or portfolio companies mentioned, referred to, or described are not representative of all investments in vehicles managed by a16z; visit https://a16z.com/investment-list/ for a full list of investments. Other important information can be found at a16z.com/disclosures. You’re receiving this newsletter since you opted in earlier; if you would like to opt out of future newsletters you may unsubscribe immediately.














The meatiest piece of thinking I’ve read on AI writing. A fun ride. Makes me wonder if the big LLMs are putting concerted effort into increasing the quality and distinctiveness of the *writing* in their default-level outputs during RLHF, or if they are primarily focused on not-wrongness, figuring that users will tune diction and style at the account/project level.
Really enjoyed this piece, guys! Thought-generation and its written expression are now the frontier for models for gaining relevance, within companies and on a wide society basis: my two-cents are that, still, passion, connecting with inner human emotions (lose the bus in the last minute) and aim to create beautiful things (maybe not useful economically) will remain in the human side. Seeing A16z writing on this pictures its relevance!