College credit for this?
What that Oklahoma Bible-citing essay controversy was really about
Social science is broken. I was glad to provoke a viral substack post this week, “A Dean Stopped Trusting Social Psychology. My Colleagues are Why,” saying yes, a great deal of social science is broken. Why should anybody believe any study these days?
The implications of the recent replication crisis for higher education are broader, deeper, and more alarming than university leaders are admitting. Millions of college degrees rest on coursework in fields that cannot vouch for what they teach, delivered through a billion-dollar credentialing infrastructure that certifies competencies the assignments inside it do not require. The rot goes all the way down to the bottom and parents should start complaining. If you are skeptical that social science is fundamentally broken, hold on as I walk you through what is already obvious.
Readers may remember the December 2025 controversy where a psychology major at the University of Oklahoma got a failing grade on an essay, written for a required online developmental psychology class, that cited the Bible to argue against the assigned study’s premise about genders. The graduate teaching assistant who failed the student said that the essay did not answer the prompt. The university subsequently suspended the grader calling her decision “arbitrary.”
The headlines were all about “woke” ideology on college campuses. I’m not going to name the student or the grader because they are not the story.
The real story is the online course, PSY 2603, Lifespan Development (also known as Developmental Psychology), a required course for all psych majors at OU and a transferable general-education option for everyone else. It sits in the Oklahoma Course Equivalency Project as PY 103 and has an equivalent at public institutions nationwide.1 Roughly 60,000 students a year take a course of this kind across the country.
In this particular course – again, taught online – students had to write a 650-word “reaction paper” to a 2014 empirical study, “Relations Among Gender Typicality, Peer Relations, and Mental Health During Early Adolescence.” The list of prompts included:
1. A discussion of why you feel the topic is important and worthy of study (or not) 2. An application of the study or results to your own experiences 3. An application of the study or results to observations about other behaviors 4. Linking the objectives or findings from the assigned article to other domains of development or other findings that we read about or discuss in class 5. A suggestion for further studies or experiments that might help researchers better understand the topic being studied 6. Alternate interpretations of the researchers’ findings 7. A discussion of how development in this domain might proceed differently at other developmental stages 8. Your own thoughts about how development proceeds in the domain being researched in the article
Everything that is wrong with American higher education is right there in the list, which basically asks students to “comment,” offering opinion or speculation, as if the paper were a social media post. I’ve put the full soul-killing assignment in the footnotes. 2
Given such a loose list, of course the student’s “reaction paper” was appropriate. Anything would have been appropriate. No brain work was required, nobody seems to have noticed. Even the American Association of University Professors and president Todd Wolfson thought the story was about a breach of academic integrity, arguing that faculty hold a professional responsibility to evaluate student work against established scientific criteria.
The twelve-year-old study in question, “Relations Among Gender Typicality, Peer Relations, and Mental Health During Early Adolescence,” a co-written paper by a University of Kentucky social scientist (and now dean) and her graduate student (listed as first author) is weak and flawed. It is cross-sectional, which means it measures gender typicality and depression at one moment and cannot say which preceded which: depressed children may report themselves as less typical, rather than atypicality producing depression. The total sample was 84 middle-school children, 34 boys and 50 girls, in a single school setting. Gender typicality was measured by the children’s own ratings on a self-report scale, with no behavioral observation, no longitudinal follow-up, no replication across schools, and no comparison across cultures or class settings. The peer-relations measure depended on the same children’s descriptions of hypothetical popular and rejected peers rather than on observed social network data. Effect sizes are modest, the moderation by gender means the headline finding holds only for boys, and the authors acknowledge the directionality cannot be established.
The only reason to assign a weak paper in a college classroom is to teach students how to identify bad methodology. In most social science programs, methodology is rarely taught and rarely taught well. Getting a degree in psychology or social work at OU does require some coursework in methodology but apparently an online course does not require students to apply that methodology in courses like PSY 2603.3
In the AI era, giving a study to ChatGPT or Claude to do a thorough analysis is the least that should be required of students in a social science course. My pro version was ruthless, identifying 20 separate weaknesses.4
Every reporter who focused on what the Bible-citing student wrote missed the real problem, which is that the assignment is misaligned with everything the course claims to teach and that the online course does not actually teach anything substantial. This is common. Public universities across the country sell a layer of the same required courses in a major, transferable across the country, under the language of critical thinking, evidence, and broad intellectual formation.
If legislators want critical thinking, students need to learn to distinguish correlation from causation, to evaluate whether the sample is adequate, to examine whether a topic such as “gender typicality” has been validly measured, to compare an older article to newer research. There is no body that ensures critical thinking is taught. With an online course, it doesn’t matter who is teaching the class, so there’s no faculty member to be held responsible. The grader was not the problem.
The December 2025 story is more worrisome than critics noticed. Everyone assumes that a psychology BA from a public flagship is evidence that the holder has been trained in the discipline. This small glimpse of one assignment in one course should be enough to make employers, graduate schools, licensing bodies, and federal lending programs start asking hard questions. Why are we giving college credit for a required course that asks twenty-year-olds to share their personal reaction to a published study and then grades them on whether they say the correct words?
The Jewell-Brown paper has been cited 166 times. Citation is supposed to mean a finding has earned its place in a canon, that it is respected and it matters. In a serious field, citations track replication and methodological strength. Here, who knows what citation tracks? A paper enters the literature by being read, gets read by being assigned, gets assigned by being cited, and the loop never returns to the question of whether the original finding had any merit.
An entire field of study is broken even as credentials in it keep getting granted. Right-wing defenders of a wronged student seem not to want to touch the possibility that the credential she was working toward wasn’t substantial. The left wants to assume the discipline has real standards and that the instructor was upholding them. I am hoping voices like Inzlicht’s will provoke a reckoning.
Every student who passes OU’s PSY 2603 class receives a real unit of transferable college credit. The Oklahoma statewide equivalency system places OU PSY 2603 inside PY 103, “Developmental Psychology (Lifespan),” alongside equivalent courses at many Oklahoma public institutions. The broader transfer system means the course would be accepted at every public university in the country.
In PSY 2113 students learn the basic statistics psychologists use to make sense of their data; in PSY 3114 students design their own small experiments, write a hypothesis, identify what could go wrong, run participants, analyze the results, write it up. By the end of the methods sequence, a competent student should be able to look at a published study and tell you whether the sample is too small, whether the measure captures what it claims to capture, whether the design supports the conclusion the authors drew. But the example of PSY 2603 shows that critiquing the methods used in an assigned study is not even an option.
“This is a weak exploratory paper that can support a classroom methods critique. It should not be used as if it gives settled knowledge about adolescent gender, peer status, or mental health. Its evidentiary reach is narrow, and several of its weaknesses are severe.
The official abstract says the study used 84 middle-school students, 34 boys and 50 girls, and examined whether gender typicality related to popularity, rejection/teasing, and mental-health outcomes. That alone should lower confidence: the article is over ten years old, based on one small sample, and aimed at exploratory association rather than robust developmental inference.
The first serious weakness is sample quality. The study used 84 students from one public middle school. The school had 1,081 students, 171 students in the relevant required health class, 94 returned signed consent forms, and 84 ended up in the analytic sample. So the final sample was about 49% of the eligible health-class group and about 8% of the school. That creates obvious consent-selection and participation-selection problems. Students whose parents returned consent forms may differ from nonparticipants in family structure, parental education, political/religious attitudes, comfort with university research, sensitivity to gender topics, school engagement, mental health, and peer status. The study cannot rule that out.
The subgroup numbers are especially weak. The entire male sample is 34 boys, divided across sixth, seventh, and eighth grade as 14, 9, and 11 boys. Yet many of the paper’s more interesting claims are boy-specific: low typicality predicted worse outcomes for boys; teasing mediated some links for boys; typicality mattered differently by gender. That means the strongest-sounding claims rest on very small subgroup cells. A regression, mediation, or interaction result with 34 boys is fragile even before considering measurement error, multiple testing, and unmeasured confounds.
The demographic profile narrows external validity. The sample was 71% White, 8% Hispanic/Latino, and 7% African American; 26% of the school qualified for free or reduced lunch; and a large share of parents had college or graduate education. The authors themselves acknowledge that the small and European American sample limits generalization. That means the study should not be treated as evidence about American adolescents generally, minority students generally, working-class students generally, or contemporary school climates generally.
The second serious weakness is the design. The study is cross-sectional and correlational. It measures typicality, peer status, teasing, and mental-health outcomes at one point in time. It cannot establish whether low gender typicality leads to teasing, whether teasing leads to mental-health symptoms, whether anxious or depressed children perceive more teasing, whether socially marginal children are rated as less typical, or whether some third factor drives the observed associations. The authors acknowledge the cross-sectional, correlational limitation, yet the article’s interpretive language still leans toward a process story.
The mediation analysis is therefore particularly weak. The paper asks whether gender-based teasing mediates the association between low gender typicality and negative mental health. Cross-sectional mediation is one of the classic danger zones in social-science inference because mediation is a causal and temporal claim: X precedes M, and M precedes Y. Methodological work on cross-sectional mediation warns that such analyses can produce substantially biased conclusions about longitudinal processes.
The third weakness is construct validity. “Gender typicality” is doing too much work. The peer-rating prompt defined typical boys as “boy-ish” and typical girls as “girl-ish,” then framed atypical boys as having qualities girls have and atypical girls as having qualities boys have. That is a blunt binary instrument. It measures conformity to perceived sex stereotypes, reputational fit, social recognizability, and possibly popularity itself. It does not cleanly measure gender identity, gender expression, psychological development, mental health vulnerability, or social-role orientation.
The wording also risks creating the construct it claims to measure. The prompt tells students that some boys are very typical and some girls are very typical, then asks them to rate peers on that basis. This can prime respondents to use stereotypes, social gossip, attractiveness, athleticism, sexuality assumptions, maturity, clothing, voice, friendship groups, or popularity as proxies. A child rated as “atypical” may be socially marginal, shy, autistic, poorer, physically immature, racially atypical in the school setting, less athletic, religiously distinct, academically unusual, or simply disliked. The study does not separate these possibilities.
The fourth weakness is that the study’s own coding reveals the datedness of its gender assumptions. In the hypothetical-peer task, popular girls were described with more “atypical” descriptions than rejected/teased girls, and the authors say this was driven by “athletic” and “independent” being treated as counter-stereotypical descriptions for girls. That is an enormous tell. A contemporary instructor should stop the class there and ask why “athletic” and “independent” are being coded as gender-atypical for girls. The coding scheme embeds a stereotype and then uses the stereotype as an analytic category.
The fifth weakness is that the hypothetical task is conceptually muddy. Students were asked about hypothetical popular and rejected/teased peers, and the authors counted typical and atypical descriptions. “Rejected” and “teased” are not the same thing. A student can be rejected without being teased, teased without being broadly rejected, feared without being liked, popular while being aggressive, or low-status while having close friends. The study blends several social phenomena into simplified categories, then infers a gender-typicality pattern from those descriptions.
The “chance” comparison in the description task is also questionable. The article treats chance as 33% because descriptions could be coded typical, atypical, or neutral. That assumes the three categories are equally likely in open-ended social description, which is not a safe assumption. Children do not generate descriptions from a uniform three-category distribution. Some traits are more verbally available than others; “popular” prompts may naturally elicit visible style, sports, attractiveness, or confidence; “rejected/teased” prompts may elicit social weakness, oddness, quietness, or vulnerability. Calling 33% “chance” makes the test look cleaner than the language task warrants.
The sixth weakness is measurement contamination. Peer ratings of popularity, likeability, and gender typicality are highly likely to be entangled. A student who is liked may be seen as more typical; a student who is popular may be seen as more typical because popularity itself defines the local norm; a student who is disliked may be retroactively described as atypical. Controlling for likeability helps only a little because popularity, likeability, and typicality are all social judgments made inside the same small peer ecology.
The seventh weakness is that “gender-based teasing” is broader than the paper’s peer-relations framing suggests. The teasing measure asked whether respondents had received discouraging comments or mockery for failing to be a typical boy or girl from teachers/coaches, mother, father, friends/siblings, other family members, neighbors, other girls, other boys, and others. That measure mixes peer teasing, family pressure, adult correction, school authority, sibling conflict, and neighborhood behavior. It is therefore weak evidence for a peer-specific mechanism.
The eighth weakness is self-report mental health. The study measured depressive symptoms, anxiety, self-esteem, and body image through questionnaires. That is common in developmental psychology, and some scales had acceptable internal consistency. Internal consistency does not solve the deeper problem: self-report symptoms collected in the same session as self-report typicality and teasing can reflect general negative affect, response style, social desirability, current mood, reading comprehension, and willingness to disclose. It is not clinical diagnosis and should not be described as if it establishes mental-health injury.
The ninth weakness is multiple testing. The study examines hypothetical descriptions, peer-rated popularity, likeability, peer-rated typicality, self-rated typicality, gender interactions, depressive symptoms, anxiety, self-esteem, body image, mediation by teasing, and alternative models. With a sample of 84, especially 34 boys, this is a high researcher-degrees-of-freedom environment. Some findings are likely to be unstable. The paper reports several p-values near conventional thresholds, and it does not appear from the accessible text to use a stringent correction for the number of comparisons.
The tenth weakness is old mediation practice. The paper uses Sobel tests in some mediation analyses. Sobel testing was common, yet even by the 2000s methodological literature had pushed researchers toward directly testing indirect effects and using bootstrap confidence intervals for mediation. With a small and non-normal indirect effect, Sobel-based inference is particularly brittle.
The eleventh weakness is confounding. The article does not adequately control for puberty status, attractiveness, body size, race, socioeconomic status, sexual orientation, disability, neurodivergence, family religiosity, school climate, classroom placement, sports participation, academic status, social-media exposure, prior bullying, baseline mental health, or parental gender attitudes. Any of those could affect perceived typicality, teasing, popularity, and mental health. A small cross-sectional study cannot untangle that web.
The twelfth weakness is school-level non-generalizability. Because the study uses one school, the school’s local peer culture may be doing most of the causal work. A school with intense athletic hierarchy, conservative gender norms, high bullying, strict dress norms, or strong clique structures might produce one pattern; a different school might produce another. The authors admit that schools may vary in how much they emphasize gender typicality, which is a major limitation because the study has no school-level variation to estimate.
The thirteenth weakness is historical obsolescence. A 2014 article based on pre-2014 data is a poor stand-alone classroom text for 2020s students unless the assignment asks students to evaluate what has changed. The study predates the current visibility of transgender and nonbinary youth, the post-2014 social-media environment, pandemic-era adolescent mental-health baselines, and the current institutional language around gender identity. The paper’s own framing refers to DSM-IV-TR “gender identity disorder,” which is a dated clinical frame for a 2020s course.
The fourteenth weakness is that it can slide from description into moralized pedagogy. The study’s basic descriptive point is plausible: children enforce gender norms, and atypical children may be teased. The problem arises when a lower-division class uses the article as a social-positioning exercise instead of a methods exercise. Students can come away with the intended moral message while learning almost nothing about sample selection, operationalization, mediation, peer-nomination design, confounding, or limits of causal inference.
The fifteenth weakness is that the paper’s strongest conclusion is also its most vulnerable one. The abstract says low gender typicality predicted more negative mental-health outcomes for boys and that some relationships were mediated by gender-based teasing. That claim rests on 34 boys in a cross-sectional design, with self-report mental-health outcomes, a broad teasing measure, and a binary typicality construct. The conclusion should be treated as hypothesis-generating. It is not a secure developmental finding.
The sixteenth weakness is that peer ratings may create dependence among observations. Students are rating one another in the same grade/school network. Peer judgments are socially interdependent: reputations circulate, friends share views, and local status hierarchies shape ratings. Averaging peer ratings can be useful, yet the unit of analysis is still embedded in one social network. With one school and tiny grade-by-gender cells, the analysis cannot separate individual traits from local network structure.
The seventeenth weakness is that it treats “typicality” as though more or less of it can be placed on a simple scale. The article itself notes that gender identity is multidimensional and that a child can be typical in some respects and atypical in others. Then the study leans heavily on continuum measures anyway. That creates a mismatch between the conceptual caveat and the statistical treatment. A child who is athletic, artistic, religious, shy, masculine in dress, feminine in friendship style, and high-achieving cannot be meaningfully located on a single “typical” line without losing much of the phenomenon.
The eighteenth weakness is interpretive asymmetry by sex. The results for girls are messier than the headline framing. Popular girls were sometimes coded as atypical because they were described as athletic and independent. Peer-rated typicality among girls was linked in the discussion to more popularity and likeability, yet also to more negative body image and anxiety. That is a tangled result pattern, not a clean developmental law.
The nineteenth weakness is the absence of strong replication inside the assigned article itself. There is one sample, one school, one time point. There is no second study, no holdout sample, no preregistered replication, no longitudinal confirmation, and no cross-cultural comparison. Later work in the area often uses much larger and more appropriate designs; for example, a 2020 study of adolescent gender-role attitudes used 4,063 students in 57 schools and multilevel latent growth analysis, which illustrates how much stronger the design can be when the question is treated seriously.
The twentieth weakness is reporting transparency by current standards. Modern APA reporting standards emphasize detailed reporting of methods, data results, analyses, interpretations, and implications. This older article reports many conventional details, yet its “alternative models” section says complete analyses and null results are available upon request. That is an outdated transparency norm. The ordinary reader cannot fully audit those models from the article text.”



"...of course the student’s “reaction paper” was appropriate. Anything would have been appropriate. No brain work was required, nobody seems to have noticed." Love this take! I haven't seen anyone else make this argument about the controversy. It was bad pedagogy and a bad essay. Bad all around.
This assignment is, in fact, gold-standard:
_X_ It offers students a "wide variety of options" for how to complete the assignment
_X_ The course is online, which means it's cheaper for admin and more convenient for students
_X_ Standards for content are lower and therefore more equitable
_X_ Word counts for such an assignment meet some arbitrary Gen Ed standard for "rigor"
_X_ The reading covers some kind of gender/race/colonial/disability aspect (often still required)
_X_ The students' own reactions and opinion must be treated as more important than any "objective" critique
This is exactly the kind of assignment that's applauded by Centers for Teaching and Learning, hall monitor faculty, and Education School pedagogy "experts."