Rigor for thee, but not for me!

I did not anticipate the first post with actual content on this new blog would be about a controversy in the social sciences and humanities, but I needed somewhere to think through and write about this issue in a more structured and long-form way than social media allows. And this is a long post, but it’s a situation aptly described by Brandolini’s law: it’s easier to create misinformation than to debunk it; or in this case, it’s easier to obscure lack of rigor behind a few short descriptions of coding practices than it is to explain why those practices are problematic at best and nonsense at worst.


The tl;dr

For those uninterested in the nitty-gritty details of bickering amongst anthropologists, here’s a quick and dirty summary. Back in June 2026, a report (the “Vanderbilt Report”) was released that attacked the social sciences and humanities for being too woke and relativist. Anthropology was identified as the most extreme of all the disciplines in the report. Many anthropologists at all levels and across many different institutions spent the days and weeks after the release of that report engaging (either positively or mostly negatively) with the Report. A few days ago, the sole anthropologist on the Vanderbilt Report team, Joseph Henrich, released his own separate sub-report with his “findings” that were used to inform the conclusions in the main report. While there is a lot to criticize about the content of the report, I instead dissect his methods to explain why his use of the Claude LLM to code and the actual codes themselves are bad, biased, and hypocritical.


The Vanderbilt Report

Some background first: In June of 2026, a report dubbed the “Vanderbilt Report” was unleashed upon the world. Commissioned by the administrators of Vanderbilt University and Washington University St. Louis, NYU philosopher Paul Boghossian led a gaggle of philosophers, historians, one sociologist, and one anthropologist—all but one from elite private academic institutions—to conduct an inquiry into the influence of (“woke”) politics and relativism—Boghossian’s pet peeve—on the social sciences and humanities. The report was laughable in its absolute lack of evidence, and when questioned by reporters covering higher education, the Report’s authors said that the individual disciplinary reports with details about the methods of investigation that led to their conclusions would probably be released at some point in the future if the sub-report authors wanted to release them, but they didn’t guarantee it.

There has been plenty of discussion and discourse, both public and private, about the contents and potential impacts of Vanderbilt Report in the ensuing months. To get a sense of how anthropologists have responded, I’ll just point to the recent Anthropology Defended issue of Anthropology News, which dropped only a few days ago. Of course, there was a broader variety of responses than these, including a few voicing support for the conclusions in the report, but as far as I can tell most of the responses were doubtful and/or incredibly critical of the Report.

A couple days ago, we finally got to lay eyes on the detailed report put together by Joseph Henrich, a biological/biocultural anthropologist in the Department of Human Evolution (which was split from the Department of Anthropology in 2009 after decades of bickering amongst bio and cultural anths there) at Harvard University. I am sure there will be plenty of people who will dig into and critique the claims he makes in the text itself (and there is plenty to criticize). But the first thing I did was look at the methods, which is where I am going to aim my criticism. Because, as I will show, not only are his methods nonsensically bad, he’s also being a hypocrite by doing the very things he whines about throughout his report and in the Vanderbilt Report itself.

Apparently, his methodological positionality (oops—a bad word!) is: rigor for thee, but not for me!

Frankly, I don’t think anyone is obligated to engage seriously with the report contents when they are based on such a flimsy foundation, so I am spending what little free time I have set aside for myself to dig deep into his methods to show why they’re so bad. I have no interest debating the polemical content he tries to obscure as based on good-faith, rigorous analysis.

Henrich’s Description of the Coding Process

The first thing I noticed when I began reading Henrich’s report is that there were basically no details about his methods, just a broad description and gestures to the Appendices. Those were not included in the full PDF of the report that was initially circulated to me by a colleague, but I found them as a separate link on the Vanderbilt website for the Report.

I’ll focus mainly on Appendix E, which purports to describe details about the methods behind his “in-depth study” (as he described it in Chronicle of Higher Education) of American Anthropologist, which is the flagship journal of the main American anthropological professional society known as the American Anthropological Assocation or the AAA. That single journal is not representative of the entire discipline, and to call what he did an “in-depth” study, as you will see, would be laughable if it weren’t such a blatant misrepresentation of what he actually did. I’m going to explain his analytical methods as I understand them from reading Appendix E before dissecting them. I’ll discuss some of the other “methods” he describes in his actual report toward the end of the post.

To try to ground his arguments about how sociocultural anthropology as an entire field has lost the plot, he:

  • focused on a single journal in the US
  • for the years 1980-2025
  • using only article abstracts (or the introductory article text when abstracts were not present)
  • after filtering out 3,713 abstracts for various reasons
  • arriving at a final sample size of 212 abstracts
  • which he then reduced to only the first 170 words for coding purposes

He used Claude, Anthropic’s LLM, to do all of the coding. No humans were involved in the coding process, nor in reviewing the output except for 10 “randomly selected” samples that he reviewed and agreed with Claude’s codes on.

He provided a code list and explained that he had Claude code in a binary way such that each domain was either present or not. The codes were:

  1. Quantitative evidence – were there numbers in the abstracts?
  2. Inferential statistics – were statistical methods mentioned in abstracts?
  3. Named systematic method – were systematic “replicable” methods named in abstracts?
  4. Advocacy stance – does the article state a “normative or political purpose” by advocating for a group, a cause, or social change?
  5. Power opposition frame – “the analysis is framed primarily around a dominant/subordinate or oppressor/oppressed structure, whether racial, colonial, gendered or class”
  6. Critical theory frame – “critical-theoretic vocabulary is used for analytic concepts: neoliberalism, biopolitics, coloniality, settler colonialism, hegemony, subalternity, assemblage, the ontological turn, epistemic violence, precarity, racial capitalism, governmentality, affect theory”

He then explained that two more codes were applied, but he chose not to discuss them at all in the report—and, curiously, he does not explain why he made that choice. Those codes were “ethnographic markers” (specific methods/terms used in ethnographic work like participant observation and fieldwork) and “reflexivity/positionality,” which, instead of offering a plain language description like all the other codes, he points down to section E.7. There, the code for reflexivity is defined as “the researcher’s own person, position, or experience [is] an OBJECT of the analysis.”

In describing how he instructed Claude to do the coding, he explains “the procedure was designed to keep the coder ignorant of anything that could cue the expected answer.” He explains he did not include the title, journal name, author, or publication year of the abstract, so the coder (again, the LLM Claude, not a human) could not tell whether the article was “a 1982 article or a 2019 one, so it could not shade its judgments toward an expected time trend.” He shuffled the batches of 45 articles coded at a time so that different eras were mixed together to try to avoid a temporal bias in the coding.

Along with the codebook, he claims he did not tell Claude the hypothesis (which, as far as I can tell, there was no hypothesis tested in his report), a description of the study, or what pattern was “expected or wanted” (wanted is an interesting choice of words, but I digress…). I’ll explain why that was not actually the case in the next section.

He had Claude’s output restricted to a JSON array, which is basically a standardized data file format for storing and transmitting data using human-readable text. He claims to have done this to make it impossible for Claude to “decline, hedge, or return prose in place of judgment.” Maybe it did do that—we can’t tell because what it absolutely did do was make it impossible to validate and have some confidence in the assigning of codes because this approach stripped the ability to figure out why codes (especially unclear or edge cases) were applied the way they were.

He then compared one archaeology journal and one biological anthropology journal for their use of quantitative methods in a purported attempt to show those fields have not experienced a decline in quantitative methods. That is only relevant insofar as he seems to think those two fields have no methodological or epistemological issues due solely to their common usage of quantitative methods.

Digging in to Bad Methods

Now that I’ve summarized his approach, I’m going to explain why it is bad and seems designed to produce a particular outcome based on a priori conclusions about the field and his own experiences within it.

1. The Sample

It boggles the mind that one would call this approach an “in-depth study” when it was based solely on the first 170 words of abstracts/article introductions from 1980-2025, all of which was coded solely by an LLM. There is nothing deep about that; in fact, it is incredibly superficial. To be blunt, it is the level of work I would expect from an undergraduate honors thesis, not a PhD-level researcher concerned about the rigor of methods in the field, especially when trying to claim broad patterns in an entire field based on that data.

An in-depth study of American Anthropologist would have at a minimum:

  • used humans to code;
  • coded entire articles;
  • not limited the start of the time frame to 1980 for a journal that began in 1899; and
  • conducted a qualitative analysis of the full articles to investigate the actual meaning and usage of his codes rather than depend solely on an LLM to quantify the output without sharing any context for human coders to interpret.

I cannot help but think that his decision to start in 1980 was intentionally meant to influence the results. It was right around the time postcolonialism made its way into anthropology, leading to the start of a disciplinary grappling with and critique of anthropology’s role in spreading European colonialism. It was also right when the Reflexive Turn was beginning to make significant changes to how anthropologists planned, carried out, and wrote about their work—just a few years before the publication of Writing Culture, which is considered a seminal text in the Reflexive Turn.

To be clear, the Reflexive Turn was and is not without controversy and introduced its own set of issues, but the introduction of reflexivity into the practice of ethnography was an important and necessary intervention into the problem of anthropologists thinking of themselves as disinterested, objective observers whose presence had no observer effects on the people they were studying. Part of what is going on here is an implicit attempt to link these changes to the proliferation of postmodernism in the academy throughout the 1980s and 1990s, and so by setting this arbitrary start date as the beginning point of his analysis, he is smuggling in an assumption that there were no debates around race, sex/gender, capitalism, violence, colonialism, and emotion before 1980. 1980 is not an appropriate “baseline” to study changes in the discipline over time through this specific journal. 1899, the beginning of the journal, is the appropriate baseline.

Henrich uses American Anthropologist as a sort of obscuring synecdoche for the entire field of sociocultural anthropology. There is no logical or rational reason to gloss that particular journal as representative of the entire subfield or, in the Vanderbilt Report itself, the entire field of anthropology. If he was actually interested in doing a rigorous “in-depth” study, in addition to using better sampling and analytical methods, he would have expanded his analysis beyond a single American journal to, at a minimum, include flagship journals for the several main areas of focus in sociocultural anthropology (e.g., medical anthropology, economic anthropology, environmental anthropology, political anthropology, media/art/music anthropology, anthropology of development, to name a few), as well as journals focused on methods, journals focused on theory, applied anthropology journals, and journals based in other countries.

And yeah, that is a monumental task. But if he is going to whinge about how the field lacks rigor, he should be rigorous himself. And hey, maybe he could have reached out to the AAA and the SfAA and marshalled their collective resources and membership to do a good faith inquiry that invited a diverse spread of folks across the discipline to assess the state of the field together. I actually think that would be a pretty awesome thing to do because I do think there are plenty of issues in the discipline for us to figure out. I just don’t think he’s identified them or used rigor to reach his conclusions about the discipline. And I think he and the Vanderbilt Report authors wanted to more tightly control the narrative, but that is speculative.

Limitations

Henrich offers three limitations to the approach in his report (bottom of pg. 42), which I want to take one-by-one.

The first limitation he describes is: “Ideally, a team of anthropologically savvy human coders, blind to the research questions, would have coded a random subset of the same articles to verify agreement.”

I agree it would have been much better for him to use human coders even though his codes and approach are bad. I do not agree they should have only coded “a random subset” nor should they have done it to “verify” the LLM. That is exactly backwards; humans should have been the primary coders, and if an LLM was used at all, it should have been used to assess human coding output, and then the human coders should have seen that assessment and offered a response to it since LLMs are notoriously good at making shit up.

Henrich could have recruited some people to do that work with him, he just chose not to. He also chose not to code it himself (maybe because he thought he would be accused of bias, so he turned to an LLM which I guess he thinks is objective, but more on that in the next section). I’m sure he would point to funding or some other logistical issues that would make this difficult, which, fair enough. But I don’t think he would have the same grace with a published article’s flimsy methods being less than ideal, especially if that article was making sweeping claims he disagreed with. My guess is he would use it as an opportunity to dismiss its findings, but we’re meant to just shrug it off when it’s his own work, I guess.

I’d also point out that he used a binary “present/not present” coding approach that intentionally and necessarily limited any context or gray areas from consideration, which human coders would pick up on and discuss together. So, this is not only an issue of it being less-than-ideal, but the very codes he created themselves point toward certain conclusions. I’ll return to that in the third subsection below.

The second limitation Henrich notes is: “we should use both LLMs and humans to code full articles, not just abstracts.”

I agree we should code full articles, not just (truncated) abstracts. I don’t personally think we “should” use LLMs to do that, but I agree humans should be involved. So why did he not do that? Time? Money? Are those valid reasons to accept his bad methods and the broad conclusions he makes from them? I don’t think so.

The third limitation he offers is: “we should apply these approaches to all of the top anthropology journals.”

I agree that all of the top journals should be included in analyses that are attempting to assess the state of an entire discipline’s epistemology. But I reject the idea that “these approaches” should be applied to anything at all. These approaches are not good and are riddled with bias despite his best attempt to pretend otherwise. (Maybe a little reflexivity would have helped with that!)

2. Claude the (Sole) Coder

Let’s talk about the eLLMephant in the room. There is currently a lot of controversy around using LLMs in general and in research activities specifically. There’s an incredibly wide range of reasons for those debates that are too much to get into this post. The widespread availability and use of LLMs is incredibly new, the software itself is riddled with issues like bias and misinformation (anthropomorphized as “hallucination”) that bleed into research activities, and unfortunately too many people treat LLMs as objective thinking machines just reporting on facts rather than as statistical word-ordering machines that lack notions of meaning or context. LLMs are biased by how they are trained, what they are trained on (they can be “poisoned” by content on the internet, for example), and how their oligarch owners and cult-following programmers fiddle with what they are allowed to output. There’s also rampant anthropomorphizing of LLMs going on, at least partly because their developers decided to give them sycophantic interfaces that create virtual echo-chambers and keep people wanting more while also leading people to form parasocial relationships with software through its constant flattering praise and encouragement, even to do harmful or illegal things.

With all that in mind, let’s look at what Henrich did with Anthropic’s Claude LLM for this report.

In a footnote at the bottom of the first page of the Appendices document, Henrich says he only used Claude on some sections of the appendix (though he does not identify which ones) and emphatically insists that “Claude did NOT assist with the main text except in explicit references to this Appendix.” I’ll take him at his word on that despite the feeling as I read through the report that some sections had a whiff of LLM writing style to it. To be fair, that could be because he is a trained academic writer and LLMs have been trained on a lot of academic writing—using em-dashes is not necessarily a sign of LLM usage as I can attest as a frequent user. Even if he did use Claude to help write the main report and this note is not truthful, it doesn’t really matter regarding his methods anyway. (I note all this because when I initially read the report and the use of Claude described in it, before I got my hands on the Appendices, I made a sardonic post on Bluesky saying this was Claude’s report with his name on it. I left it up and responded to it with critiques of the methods, and I don’t want to pretend I didn’t say that publicly, so it is what it is. Mea culpa.)

The first issue is that Henrich used Claude to conduct all of the coding. No humans coded anything. Out of the 212 truncated abstracts he had Claude code, he “randomly selected” (and did not explain how that process worked) 10 of them, about which he reports he “fully agreed with the LLM’s decisions” (pg. 42 of Henrich’s report—note the anthropomorphizing language “decision”). To put that in quantified language Henrich might better appreciate, he only checked 21% of the coded abstracts. Why not check all of them? He doesn’t say. I assume he would say that he simply lacked the time, but to me that would be a cop-out he wouldn’t (and shouldn’t) accept from other people taking shortcuts in their methods.

As described in the previous subsection, he meekly admits this dependence on Claude to do all the coding is a limitation in one short sentence without a discussion about why. So, I’ll do it for him by closely examining his approach as described in Appendix E and dissecting it to expose the various ways it enables bias to be smuggled in, even as Henrich implies his approach stripped the possibility of bias from the process.

Before digging into the specifics of his prompting and coding variables, I want to note that Henrich says at the end of Appendix E section 2 that when he discusses Claude as a coder, it “represents a ‘fresh’ Claude Opus 5.0.” That is the extent to which he discusses how he began the coding process with Claude, and it is not as helpful as he might have thought. What does he mean by “fresh” here? A new chat thread within his main account, which would have access to every other chat thread he’s used Claude for, including any conversations he had while planning the project, collecting data, or even just brainstorming about the project? Was it a completely new account with a different email address from a main account that was totally detached from any previous use of Claude as well as any other accounts he owns in which he did any work around this project? He doesn’t say, and so it is impossible to evaluate based on the provided information just how “fresh” his coding instance of Claude was.

He also does not provide the entire conversation thread, just one prompt and the codes, nor does he provide (as far as I can find) the actual raw pre-coded data itself for others to inspect. He does not describe how many rounds of prompting or coding he did, or any tweaks he made to prompts in the process. Was this the first and only prompt he used, or did he have to “prompt engineer” to get it to produce the output he wanted or thought was useful? A detailed description of his entire process would go a long way toward making these things more understandable and assessable (and, dare I say, replicable, to refer to one of his own codes used to demonstrate the lack of rigor in sociocultural anthropology). Frankly, given the criticisms he is making, it comes across as obscurant and like he is hiding something by leaving out these kinds of details, which makes me suspicious that he is trying to make his methods seem objective so he can more confidently share his sweeping claims about the discipline. I’m happy to be shown wrong should he release all of these details and data.

All right, let’s (finally) get into the details about his Claude coding process. In Appendix E section 7, Henrich provides the prompt he used for what he called “anthropology methods content analysis”:

You are a research assistant coding journal-article abstracts for a study of how research methods and epistemic framing in anthropology have changed since 1980. Code ONLY what the abstract itself states or clearly implies. Do not use outside knowledge about the article, the author, or the journal. Do not infer from the topic what the methods “probably” were.

The first thing to note about this prompt is that Henrich said multiple times that Claude was not provided a hypothesis or description of the study, but the prompt begins by explaining the description of the study (if not the hypothesis which, again, he never gives us in the report—what, exactly, was he “testing” if there was a hypothesis?). The prompt explicitly tells Claude that Henrich is investigating temporal changes in methods and epistemology over time since 1980, and thus he has given Claude information about what he is looking for. If he wanted to truly blind the coding, he would not have given this information—and he certainly should not have claimed that he did not provide it when he clearly did.

In the second sentence, he tells Claude to code only what the abstract states or “clearly implies.” That usage of “clearly implies” opens space for interpretive work to occur, but again the way he had Claude produce its output in JSON format with binary codes prevents him from checking that kind of thing. This is why human coders are important in this process—it’s not that “clearly implies” is inherently a bad thing to consider while coding, but it necessarily means there is not always a binary “yes present/no not present” choice, so some things must be inferred based on the linguistic context present in the abstract, context which is often lost on LLMs. As an LLM, Claude operates using theory-free inference and thus auto-generates text that is a best approximation of what a human would say based on its training data without regard for why humans might say that (and in that way it also cannot generate things that are entirely novel or unexpected).

At any rate, Henrich tries to have it both ways: he claims there’s no inferring or hedging allowed but his directions allow for it anyway. And because of his approach to the output generated by Claude, the process of inferring or hedging becomes invisible and is outputted instead as “objective” yes/no information even when its possible some cases are not clearly binary and require some nuanced interpretation of the context by humans.

The last two sentences in the prompt instructing Claude not to use outside knowledge or infer methods from the topic are noble attempts to limit bias into the analysis, but are ultimately not as useful as Henrich makes it seem. First, it’s not really possible to have an LLM “not use outside knowledge” because LLMs do not cordon off information like that. They are not thinking/reasoning machines but word prediction machines. Everything Claude coded for Henrich was influenced by everything it has been trained on, so it does not need to “look up” details about, for example, “critical medical anthropology” being used more frequently after the term was introduced in the early 1990s in order to make predictions about words and auto-generate textual outputs for Henrich. In that sense, abstracts have a sort of built-in metadata that don’t require Claude to know an article’s title, author, or year to bias its output because vocabulary and writing style are noticeably different over time.

Similarly, prompting Claude to not infer what the methods in abstracts “probably” were from the topic of the abstract isn’t as useful as he seems to think. Again, I think this is a noble attempt to reduce bias and avoid shortcutting in coding based on words mentioned rather than a literal description of methods. But these two things—the topic and the methods—are not so easily teased apart in the way he wants it to be. Claude does not engage in inferential reading comprehension because it is a statistical word-predicting auto-text-generator program, not an interpretation/meaning-generating machine. For example, for Claude to code “self-accounts from 15 women” (the example Henrich used to train Claude to use in the Quantitative Evidence code) as quantitative requires it to draw on the same training processes that will lead it to predict that “self-accounts” tend to be qualitative. He can force its output into one or the other, but that does not mean that Claude applied a code unproblematically or without making that kind of predictive connection in other cases but in the opposite way than he instructed. If Claude was trained to associate sociocultural/humanistic topics with certain kinds of methods (read: qualitative), Henrich cannot ask it to not use that in its coding process because it’s baked into Claude’s modeling process. All he’s actually accomplished with this prompt is to hide that process and information behind an illusory objective output that he forced it to produce.

Just to re-iterate, this coding approach as well as his instruction in the prompt to “not add commentary before or after the JSON” (Instrument 1 in Appendix E section 7) prohibits investigation into potential false positives and false negatives in Claude’s output, and thus no one is able to check the output for that kind of thing. Having humans be at the center of this process would have been much more transparent and, therefore, reliable.

3. Dissecting the Codes’ Constructions

Time to examine the specifics of the codes Henrich constructed for Claude to use. In the interest of not making this already-lengthy post even longer, I’m going to just point out a couple of issues I noticed with the first three of his codes before getting more detailed about the remaining three codes.

As a reminder, the first three codes were “quantitative evidence,” “inferential statistics,” and “named systematic method.” The codes are framed as neutral but, in fact, tilt toward his a priori conclusions. For each of these codes, he provided a definition and how to apply the binary code. As I explained in the previous section, the way he set up these codes inherently “unblinds” Claude in some ways. Methods often point to subfields, years, and sometimes even authors. He also conflates “systematic” with “replicable,” which should be two separate codes.

Plenty of methods are done systematically, including participant observation, fieldwork, and “unqualified” interviews (meaning that in the abstract it did not explain what type of interviewing was done, which is why coding entire articles would be useful as that information is probably provided in the full text). But he instructed Claude to count those as “not systematic.” This built his conclusion into the code itself by telling Claude the most basic descriptions of methods commonly found in short article abstracts should not be counted as systematic. That does not, in fact, “measure” systematicity in methods of these articles; it only shows the specificity of descriptions of the researchers’ methods in the truncated abstracts.

What he’s doing here is a sleight of hand. Quantitative evidence and inferential (rather than descriptive) statistics are taken, prima facie, as evidence of rigor and are coded so that he can point to the apparent decrease in quantitative methods in American Anthropologist as a general decline in rigor across the entire field of sociocultural anthropology. It is very clear from this that he believes quantitative work is automatically and inherently better, more rigorous, and more objective than qualitative work. That is simply false, and the evidence is his very report. He has shown that you can devise quantitative work that is utterly nonsensical and lacks rigor, but because people assume quantitative work is inherently more objective, he can more easily get away with passing off his biases as rigorous scholarship.

In an 1899 article in American Anthropologist, the “Father of American Anthropology” Franz Boas wrote about criticisms being levied at that time toward physical anthropologists and their quantification practices around bones and skulls: “I believe the tendency of developing a cast-iron system of measurements, to be applied to all problems of physical anthropology, is a movement in the wrong direction. Measurements must be selected in accordance with the problem that we are trying to investigate” (103).

I think this lesson holds true across the entire discipline (indeed, across all research disciplines) up to the present day. There is a reason we choose to use some methods and not others to answer research questions. Quantification is not always a necessary or even interesting way to approach research (as, I would argue, Henrich has clearly demonstrated in his report). The argument Henrich makes, that a decrease in quantitative methods in the abstracts of one single journal marks a decrease in rigor across an entire discipline, is nonsense. Like any standardized research methods that are done well, qualitative work can be done with rigor. And like any methods done poorly, quantitative work—like what Henrich does in this report—can lack rigor. The presence, absence, or frequency of qualitative, quantitative, or mixed-methods tells us nothing about the rigor of the work itself, especially without looking past the first 170 words of abstracts or introductions.

It is also worth noting that, even though he says he is not trying to claim a causal link between the drop in frequency of quantitative methods and the rise of “postmodernism” in the discipline, the graph that juxtaposes them invites such a conclusion anyway. Appendix E section 4 notes that quantitative evidence fell from 42% to 13% between the first two “eras,” accounting for the vast majority of its decline in these abstracts. But during that same “era,” advocacy only rose from 3% to 13% and critical theory from 0 to 11%. So clearly something else had to have influenced the shift in quantification than a slight increase in advocacy and the use of critical theory keywords; yet, his report still makes it sound as if those two things are deeply intertwined.

The final three codes are where I want to focus on some details in the codes themselves because those are really where he gives away the game in the coding process. I’m going to include screenshots from the report itself rather than copy-pasting and formatting all his coding instructions.

Advocacy Stance

This is how Henrich instructed Claude to code for advocacy stance:

It is a peculiar choice, in an ostensible analysis of truncated abstracts, to instruct Claude to consider the “ARTICLE ITSELF.” It did not code the article, unless this wording then prompted it to silently investigate the article if it was included in its training data. Which we don’t know because of how Henrich told it to generate its output.

At any rate, again there is an issue with assuming Claude can distinguish between “calls for” advocacy and “merely describing or studying” advocacy, because Claude is not a meaning-generating machine but a predictive text machine. Henrich has tried to make Claude infer authorial intent here despite claiming to not want it to do so.

Is there a meaningful difference between taking up a “normative” versus a “political” purpose? If so, what is it, and why does he not explain it? Why blend the two? If he thinks these are synonymous, it seems hypocritical of him and the other authors of the Vanderbilt Report to make a normative call for a “disinterested” stance in social science and humanistic research while also saying they have no political motives. It’s either normative=political or it’s not. His codes run counter to the claims he and the Report authors want to make.

The biggest problem beyond the code as structured is the purpose behind it. Henrich makes an Olympic-level logical leap from the presence of advocacy or calls for advocacy in article abstracts in one journal to the claim that calls for advocacy inherently contaminate knowledge/epistemology and necessarily lead to less rigorous work discipline-wide. That is an unsupported conclusion and, in fact, cannot be a conclusion he arrives at based on coding for the presence or absence of calls for advocacy in abstracts in one journal. It requires a detailed coding of the full text of articles, their detailed methods descriptions and theoretical approaches used, the conclusions drawn, and how their advocacy fits in. Did they investigate something and discover some great harm being enacted, and in the conclusion of the article say “people should stop allowing this to happen” and that was the extent of their advocacy? Is that the same thing as someone doing an engaged ethnography where a community has approached them and asked them for help on a particular issue, which is how the project was shaped and carried out? And does that approach mean the knowledge produced through that engaged work is inherently poisoned by wokeism and relativism? These are inherently things this approach cannot uncover or address in any meaningful way because it requires qualitative investigation.

Power Opposition Frame

The entirety of the code for this is:

One single sentence with four keywords to, I guess, narrow the frame? If the article was about the oppression of people with disabilities, I guess that is left out? Or does he just assume Claude will figure that out despite not offering that keyword in the code itself? How does Claude determine what counts as “dominant/subordinate” or “oppressor/oppressed”? He gives no specific rule for applying the binary code like the other previous codes. This is baffling.

There is also waffling language in the code: what does “primarily” mean? How is that quantified or identified? An abstract that mentions sexism or racism in addition to its main arguments or findings will be coded as “yes” the same way an article entirely organized around an analysis of sexism or racism would. Because, again, Claude is just a statistical guessing machine and does not interpret the meaning—especially with a code like this that offers zero instruction for its application.

The same logical leap as the advocacy code is here, too. The presence of analysis of power structures does not necessarily lead to a lack of rigor. That is undemonstrated by this code and requires qualitative analysis.

Critical Theory

The code reads:

This code similarly has no instructions for application and instead prompts Claude to code “yes” for the presence of several keywords from various theoretical approaches generally grouped here under “Critical Theory,” which itself is a very loose collection of some dramatically different theoretical schools. The lack of application guidelines and only coding for the presence of these keywords does not actually distinguish between use or criticism in the article, whether it is mentioned incidentally or is central to the entire article, and cannot be used to code negative examples. As far as I can tell, this code should have just been named “confirmation bias.”

And by the way, you’ll never believe this, but these schools of thought began working their way into anthropology in the 1970s and really began to take center stage in the discipline’s theoretical repertoire after the Reflexive Turn in the late 1980s. And, surprise surprise!, Henrich’s code identifies a dramatic rise in the use of a specific set of critical theory terms in American Anthropologist abstracts as more people began using the terms in their work and analysis. There’s also a problem that he does not separate out these keywords in a way that we can actually tell which ones are more frequent. He just unloads a mass of words for Claude to flag, so of course there was such a dramatic rise given he gave it 14 keywords compared to the other codes—notice the much less steep rise in frequency? That dramatic rise is at least partly an artifact of how he set up the code.

Now, I actually do think this is indicative of an real trend (not sure to that degree), and I do think it is a discipline-wide shift not limited to American Anthropologist. But that’s not because of some nefarious erasure of rigor in the discipline or an intentional and forced replacement of previous approaches. It is literally just a measurement of the diffusion of theory and anthropologists deciding that some new terms are more useful in their work than other older ones. A good control against this would be some other more neutral theoretical terms like “culture.” But I don’t really get the sense Henrich was interested in developing good controls.

Like any field, anthropology undergoes paradigm shifts. The discipline has had several in the past 100+ years: unilineal evolutionism, functionalism, structuralism, symbolic-interpretive, historical particularism, cultural ecology, cultural materialism, poststructuralism, and on and on. My guess is that Henrich and I would probably be in agreement that poststructuralism is becoming a tired paradigm in anthropology and we should start trying to find new ways to make sense of the world that aren’t grounded in mid-20th Century French philosophy. But that doesn’t mean I think concepts from poststructuralism are no longer useful, just as I don’t think the work of Mary Douglas is no longer useful even though we have moved on from the paradigm of symbolic-interpretive anthropology.

And once again, this code is used to make an unfounded logical leap from “presence of terms” to “no rigor” without any actual supporting evidence or qualitative analysis.

Weaknesses

At the very end of the Appendices, Henrich acknowledges the weaknesses of both the power and critical theory codes in this way:

CRITTHEO is, in practice, a vocabulary checklist: it names fourteen terms and supplies no rule for the case where such a term appears incidentally rather than as the analytic frame, and no example of a 0.

POWER is a single sentence with no negative rule and no example at all. Both are therefore more exposed to coder drift than the methods codes, and both are the codes on which an abstract’s topic is most likely to be confused with its frame — the very error the ADVOC rule was written to prevent.

So instead of “correcting” this limitation before coding, he ran it anyway knowing that the structure of these two codes—which form much of the basis of his complaints about the discipline—would necessarily be problematic and inaccurate. Despite this very glaring problem and its impact on the results, he nonetheless carried on using these “data” to make sweeping claims about the entire discipline.

Reflexivity

Before moving on from his codes, I think it is important to look at a code that he used but did not discuss in the report or Appendices and does not explain why. And I think this is actually one of the most important of the codes (as bad as they all are) because it is so fundamental to contemporary anthropological ethnography.

 

Again, you may be shocked to learn that after the <<Reflexive>> Turn, it became standard practice in anthropology to acknowledge our positionality within our fieldwork. The purpose of this is to try to lay bare our biases so others can see if there’s things we might have missed or misinterpreted. The presence of reflexive language, again, cannot tell us anything about the rigor of the work.

Henrich chose not to share the results from this code and did not explain why, so I am left to speculate. My best guess is that the code did not show a rise or even much of a presence of reflexive language in the truncated abstracts of American Anthropology. If that’s the case, I think it’s because, in my experience at least, reflexive commentary is typically buried in the text of the article itself—in the methods section, the discussion section, or sometimes even embedded in the analysis—rather than shared in the abstract. Having a flat line at the bottom of his chart would not have lent itself to the story he wanted to tell about the discipline, so it’s absent. Or maybe it was another reason all together. But in the interest of transparency and adhering to his own standards of rigor, he should have shared it regardless of what it found.

That said, I would absolutely expect that a rigorous analysis of the anthropological literature would find a dramatic rise in reflexive and positioning language post-Reflexive Turn because of exactly what I said above about it becoming a writing norm in the discipline. In other words, such language would, to me, be a sign of rigor, NOT a sign of its decline.

Limitations

A few other limitations I want to briefly note about this approach:

  • If a hypothesis was tested, it was never shared.
  • There seems to have been only one coding pass by one non-human coder. Typically you code multiple times; if this was done, it was not described or explained.
  • No validation – Henrich’s “random sampling” of 10 abstracts that he “fully agreed” with the conclusions about does not a validation make!
  • The choice to truncate abstracts/introductions when no abstracts were available was vaguely explained and the choice to limit it to 170 words had no detailed justification and no explanation of how he checked the sensitivity of that choice.
  • He lumped linguistic anthropology in with sociocultural anthropology even though they are two different fields and many linguistic anthropologists use very different methods than sociocultural anthropologists. There is certainly overlap, but this approach necessarily erases the specific norms of linguistic anthropology while he keeps those of biological anthropology and archaeology separate, even though there is also overlap between those fields and sociocultural anthropology.
  • The “subfield” rule, which I didn’t cover here but was the second coding instrument he used in which titles were supplied to Claude, has the problem of grouping things under biological anthropology or archaeology that could also be considered sociocultural, for example “human behavioral ecology,” “genetics” (lots of sociocultural medical anthropology on this topic), and “ethnohistory grounded in excavated evidence” (did the researcher excavate it themselves, or did they use evidence excavated by others in their ethnohistorical work?). These are not so simple to divide up as he makes it seem, and this in fact moves some work that might have had quantitative methods out of the domain of sociocultural anthropology and situates them in other subfields instead. The direction he gives to Claude to “judge by the QUESTION the article asks, not the technique it uses” only contains examples about confusing archaeology with biological anthropology, confusing those fields with sociocultural anthropology apparently wasn’t a concern he had, which is strange considering he is a biocultural anthropologist.

Do What I Say, Not What I Do! Hypocrisy as Method in Henrich’s Report

Henrich’s report mentions some of the other sources of “data” he used to draw his conclusions. I don’t want to go into detail about all of them, such as using seven AAA presidential speeches and one keynote, all post-2013, to make claims about the decline of what he sees as good norms and values of the entire discipline, or using a handful of elite-level universities and their internal battles as representatives of discipline-wide concerns or experiences (I have personally never experienced issues in either of the four-field departments I came up in, and in my current department as an applied anthropologist we are desperately trying to get hire more bio anth people but…budget cuts and all that fun stuff).

But I think when we take a moment to reflect on what he is basing his claims on, we can pretty clearly see that he has been incredibly hypocritical in this report. Much of his criticism of the discipline as captured by liberal politics and wokeism is grounded in what he sees as its lack of rigor and its mishandling of data, especially around collection and analysis. But he sure didn’t let that stop him from throwing rigor out the window when it came to this report!

“Interviews”

Henrich states in several places in the report that he “interviewed” various people, but he never explains how that was done. All he does is tell us that he chatted with people here and there about various things he thought was relevant to this report. Where’s the rigor? Was there a standardized interview guide? What was the recruitment strategy? How many people did he talk to because they approached him versus he approached them? How many people did he ask to participate that declined? Did he undergo IRB approval for ethical human-subjects research, and if so what is the approval number? If not, why not? What was his total sample size and what is the justification? How did he ensure he recruited more of a representative sample than just chatting with his friends and acquaintances? Who, if anyone, did he talk to outside of his elite academic circles? What subfields were they part of? Did he get a mostly even split between the subfields? Did he talk only to academic anthropologists or did he also interview practicing anthropologists outside academia who keep up with and publish their own academic scholarship? These are the kinds of things a rigorous approach to interviewing would consider and do.

In writing up, he kept total anonymity for all the people he “interviewed.” Anonymity for participants is standard in ethnographic work and not itself a problem, but if he’s not going to actually share details about things he heard, it is not helpful or in good faith to drop in these little comments like he heard “several disturbing stories” because all that does bias the reader in favor of his a priori conclusions about the field. Why should we believe that claim?

On that note, why did he choose to omit the “several disturbing stories” rather than change identifying details to maintain anonymity so he could convey their important lessons? If there were actually several stories, he could have pretty easily blended details to share the main gist of what they conveyed. There is a rigorous process for that, by the way.

Obviously based on what Henrich wrote in the report, there was no attempt to make these “interviews” systematic or rigorous. Instead, what he has done is exactly what he accuses sociocultural anthropologists of doing: selective reporting on things that fit his preconceived conclusions. I think it also demonstrates his discontent for qualitative research and that his default position is that qualitative work is incapable of being rigorous. And if he doesn’t think that, it’s impossible to tell from this report.

“Auto-ethnographic Musings”

I want to wind down this critique of Henrich’s methods by briefly discussing what Jon Marks referred to as Henrich’s “auto-ethnographic musings” in a Facebook post (which I would link to but it is not publicly available since Marks had to make his account private after being attacked by right-wing freaks last year). But I think it is such an apt description of what this report actually is. It’s also a perfect example of how hypocritical Henrich has been in the report.

Henrich spends a few pages railing against autoethnography, which he defines as when “the researchers themselves or one of their intimates (their own father, mother or grandparent) take the central role, usually in a narrative of struggle and suffering as they confront the challenges of life, with heavy emphasis on the experience of racism, sexism, capitalist exploitation and othering” (pg. 37-38). Talk about a loaded definition! It’s almost as if he wants the reader to begin from the perspective that all autoethnography is done this way (and that all autoethnography is therefore nonsense) rather than that it is a particular method that can be flexibly deployed in myriad ways.

Carolyn Ellis and colleagues (2011) offer a much more neutral and useful definition of autoethnography: “an approach to research and writing that seeks to describe and systematically analyze (graphy) personal experience (auto) in order to understand cultural experience (ethno)” (pg. 273). They immediately go on to say, “This approach challenges canonical ways of doing research and representing others and treats research as a political, socially-just and socially-conscious act” (pg. 273, internal citations removed).

I hope this illustrates how incredibly biased Henrich’s description is. It should be uncontroversial to note that research is political (often in ways that aren’t immediately obvious, like what research is selected for funding or where and when people are allowed to conduct research—the IRB was created through political processes!). And speaking of the IRB, research being just and aware of social impacts are part of the three core principles of research ethics identified in the Belmont Report, which is one of many documents that guides ethics in human-subjects research in the US. To act as if autoethnography is doing something radical by having that orientation is nonsense.

I have used autoethnography as a method, and not for the reason Henrich argues it is used for. As part of my dissertation work at an anal cancer prevention clinic, I underwent the same High Resolution Anoscopy procedure that the patients at the clinic underwent. I did it because I wanted to have an embodied experience of the things patients explained to me. And you know what? It helped me understand things so much more clearly, and it also enabled me to build rapport with patients much more easily, and thus they opened up to me about things more once they knew I had had the experience as well. But I guess that was not good methodology in Henrich’s view because…all autoethnography bad!

For all his complaining about autoethnography, his report is utterly dependent on it. He begins the report with a story about his arrival at UCLA in 1993, along with an autoethnographic accounting of his intellectual journey from cultural to biological anthropology. His analysis is, to quote a critic of autoethnography, one in which he “take[s] a central role…in a narrative of struggle and suffering as [he] confronts the challenges of life” (Henrich pg. 37-38). It’s just that rather than emphasizing “the experience of racism, sexism, capitalist exploitation and othering” (pg. 37-38), he instead emphasizes the experiences of ideological pressure, professional intimidation, and the marginalization of dissenters in the discipline (of which he is apparently one). He bases his understanding of the biological/cultural split in some departments almost entirely on his own experiences in two different departments at elite private schools. There really is no attempt to be systematic and rigorous in his analysis of those events in a broader, more general way than what he and his friends experienced.

Throughout the report, he explicitly states his own positionality as exactly the kind of credence-lending thing he rebukes when others do it. “My immersion in these other disciplines, which includes earning tenure and promotion to full professor, gives me a unique perspective on Anthropology” (pg. 3) may as well have been phrased, “as an anthropologist myself…” This is precisely the kind of work autoethnography and reflexive positioning does, but I guess it’s only verboten when it is Others positioning themselves in relation to structures of oppression. In his case, it’s just adding context and is totally apolitical!

At the top of pg. 39, Henrich quotes Susan Pickard’s critique of autoethnography:

What distinguishes self-expression from scholarship? What are the limits of subjectivity as data? What about distortion, selective seeing, the dangers of solipsism? The speakers were apparently unbothered by these elementary concerns, preempting scrutiny or critique by deploying the concept of “positionality.”

Indeed! What distinguishes Henrich’s self-expression from his analysis of the discipline of anthropology? What are the limits of Henrich’s subjective experiences in particular places at particular times as data? What about distortion, selective seeing, the dangers of solipsism around Henrich using his own experiences as generalizable or not seeking out talking to people beyond his own circles? Henrich, too, seems unbothered by these “elementary concerns” and seems to be using them to preempt scrutiny or critique by using his own positionality to claim and communicate authority. Henrich doesn’t treat his own positionality as something that might bias his experiences, perceptions, and explanations of what happened at Emory or Harvard, he only invokes it as a cloak of authority so the reader is influenced to take his accounts at face value as objective. After all, who are we to question his experiences?

The main hypocrisy here is that Henrich isn’t really interested in having a good faith conversation about the usefulness or disadvantages of using authoethnography (he can’t even define it fairly), or exploring how it might be more standardized in anthropology. No, he explicitly dismisses it as an invalid method of inquiry—especially when practiced by or about groups he feels are too woke or too political—while simultaneously using his own experiences as a way to impugn not just autoethnography but the entire field of sociocultural anthropology.

Wrapping Up

On September 25, 2026, the Chronicle of Higher Education published an article containing an interview with Henrich. In it, he shares a few insights that I think are worth commenting on:

  • When asked if he had any trepidation about the commission, he said he knew some people would be unhappy but he participated because he really cares about the field and wants to improve its “epistemological health.” He immediately follows this by saying, “I think just being straight up about it and trying to study the problem and come to an understanding of it, and then write it up, was really what I tried to do.” While I appreciate his expression of care for the discipline, he utterly failed to achieve his goal and has contributed to harming it instead. He was neither “straight up about it” nor did he actually try to “study the problem” (whatever that might be).
  • He said he decided to release the report because he “want[s] to have a discussion about how we can do anthropology and other parts of the humanistic and social sciences better…so I wanted to kind of generate that discussion.” This is the full elitist academic ego on display: as if these are not already constant conversations going on in the discipline and he is the person to “generate” it for us. He is more than welcome to join us, but that would require him actually connecting with us instead of holding us at arm’s length like he picked up a bag of stinky dog shit at a park.
  • When asked if he tried to interview anyone on “the other side” (of what, I’m not entirely sure, I guess of this “debate” if we could even call it that?), his response was: “Well, I mean, I feel like I did consult with them in the sense that I read presidential speeches since 2013, and I made an in-depth study of American Anthropologist.” Hear that, folks? Reading speeches and having Claude do all the nonsense I described above is the same as “consulting” and interviewing people you disagree with!
  • Someone else was supposed to be his “partner in crime” but, he says, “the people we queried recognized that there were issues and concerns but did not want to participate because of the potential fallout.” Oh, really? Name them. Or encourage them to come forward and explain why they made the decision. Was it actually because of the “potential fallout”? Was it because as they learned about what this commission was doing they did not want to participate in a bad-faith political hitjob? How are we to know?! Let’s just take his word for it. And where, exactly, were these people located within the academy? Frankly, if I had been asked, I might have joined it just to be a dissenting voice and try to keep them honest. But, you know, it’s only “the other side” who suppresses dissent. It’s not like the people on this commission would want to exclude people from their work who disagreed with their methods and arguments…right?
  • He says he wishes he had more contact with the folks over in the Department of Anthropology at Harvard and he thinks it’s a loss that the department split happened. The editor added a comment that they reached out to the chair of Harvard’s anthropology department, who said he would welcome discussion with Henrich and the others involved with the Vanderbilt Report. So why didn’t Henrich just send the chair an email and ask to pop in to a faculty meeting some time? Something tells me there’s more afoot here—either Henrich simply did not reach out to them about any of this and he’s trying to save face, or he intentionally does not have a relationship with that department for some reason. I can only imagine what people like Arthur Kleinman and Byron Good would have to say about all this.
  • He was asked about the cancellation of a panel about biological sex before the 2023 AAA conference, which gained national attention. He, like many others, misrepresents what actually occurred there. First, it never should have made it past peer review because the panel description was making claims not in line with settled science in biological anthropology. But more than that, the intent of the panel was to try to use biology to diminish the dignity and rights of trans people and abolish the use of “gender” in the discipline—the use of “gender critical” is a dog whistle used by anti-trans activists. And the panel was organized by right-wing anti-trans activists, almost none of whom were anthropologists, as an attempt to legitimize their movement start shit within the discipline. The panel was never intended to be a good faith discussion or debate about sex and/or gender. This would be like people getting pissed because a panel by “evolution critical” people made it past peer review at a biology conference and was cancelled after people noticed it was a dog whistle for creationism.
  • Henrich was asked about the inclusion in the main Vanderbilt Report of a study from 2023 about women and hunting in foraging societies, which was later corrected for methodological issues. In the Report, it was used as one of the few field-specific attacks on how gender ideology has run amok in anthropology despite the fact that that whole thing is actually an example of science doing working correctly. The interviewer asked why this study was in the main Report but not in his own report, to which Henrich replied that he removed it from the published version of his report because the report was “getting too long and I decided that I didn’t like that example as much.” Well, isn’t that convenient! One of the handful of examples the main Report depends on is one that doesn’t actually support its conclusions, so he just handwaves it away for exactly that reason. That is, definitionally, cherry picking.

The interview concludes with a question about the response that Henrich predicted versus what he has experienced. Just before the conclusion of his report, on pg. 61, he predicted:

Many high-profile anthropologists in the field will not only demonstrate that they disagree with me — by offering alternative lines of evidence or reanalyzing the data I present (I welcome this) — but also dismissing the evidence I’ve presented by questioning my character, morality, and motivations. In doing so, they will train another generation of junior scholars and graduate students in our field to keep their concerns to themselves and stick to the party line.

He admits to the interviewer that, “So far, so good. I’m wrong. I have not received such a reaction. I’ve gotten a few kind words from a small group of anthropologists.”

I expect that in the coming days and weeks, as his report circulates and people take time to read through it, he might get some of what he predicted. While I’m not a “high-profile” anthropologist (the elitism is strong in that one!—oops, did I question his character?), I actually could not give less of a shit about Henrich’s character, morality, or motivations. Good people do bad things or make poor choices all the time. I don’t know him and have nothing substantive to say about his character or motivations. I suspect we’d probably get along fine and be able to find a lot of common ground if we were to have an honest, good-faith conversation.

In truth, I used to love teaching his W.E.I.R.D. material in my intro classes because I found it granted a lot of students an easy way to begin understanding culture, ethnocentrism, and cultural relativism. I stopped using it a couple of years ago because I switched to a textbook and removed a lot of supplementary readings, not because of any of this report stuff. The only feeling I really have about him is that this report makes me sad that he thinks relativism should have no place in anthropology (which is one of the points of the Vanderbilt Report, a la Boghossian’s pet peeve that I mentioned up top). This report could have been completely anonymized with his “findings,” and I would have had the exact same issues with its “methods” and the hypocrisy that forms the foundation of the claims made in the report against sociocultural anthropology.

There is plenty for us to work on improving in the discipline. We’re already doing that, and I for one would welcome people with critical views as long as they’re offered in good faith and an openness to listen to others in the field. Unfortunately, Henrich’s report and the Vanderbilt Report it contributed to will do absolutely nothing to aid those efforts. Instead, regardless of what the authors claim or hope the reports do, they have handed axes to administrators looking to chop down programs and kill them off. It’s just an even greater shame that axe was made of fantasy, fallacy, and bad coding.