#75: Against Nostalgia (Shrek is 25)
It is still an AI newsletter, just skip the intro
Yes, Shrek is 25. We were watching it the other day and were shocked at:
How well the animation/CGI held up (I’m looking at you, Death On The Nile)
How well the story and message held up (also Shrek 2, by the way).
Of course, we grew up with it. It’s basically our comfort film.1
Perhaps ‘comfort’ is the operative word here.
You know, you start thinking: the world is going to shit; how come I can’t go back to that time where there was internet, but no social media? Or playing outside was safe? Pre-9/11, post Cold War. LOTR, Episode III, Donnie Darko.2 No wonder the Matrix said that peak civilisation was around 1999-2000 lmao.
Either way,
I argue that nostalgia blinds us to our current problems. It allows us to seek comfort in the familiar, instead of facing the present.
It is nice to think of 90s’ malls being full and being carefree and whatnot—and hopefully you do. It means that you had a good childhood.
You don’t miss the 90s. You miss being young.
It wasn’t the same for everyone. It was a pretty homophobic and racist time by today’s—ideal—standard: just think of the slang. For many the 90s symbolise long periods of violence and war.
Or, to centre me a bit more, the 00s: where anorexia, objectification, and body shaming were pretty brutal. That marginalising slang was still there, too.3 And, FWIW, I did not grow up in the US. I’m guessing it was the same there.
This is very blatantly exploited by the media and corporations: you now get reboots every five minutes of beloved classics. Most of them miss rather than hit (cough Star Wars cough) and yet we still watch them/play them/consume them. It is 50% attracting new demographics (if you have kids they are probably way more into the new Star Wars than you are), and 50% the New Slurm.
Enjoying it is not wrong, of course. That’s not my point.
My point is that, unfortunately, today, we are definitely heading to some sort of technofascist state, in no small part powered by irresponsible development and application of scientific research. We don’t think about it that much. Like, we all know it, but then what?
I don’t think it is wise to bury our heads in the sand and glue ourselves to our disneyfied screens and pretend that there is nothing that we can do (I mean, do you have Instagram? How often do you use ChatGPT? When was the last time your research had misuse mitigations in place?)
The truth is that there will always be ‘good old days’ because we seek reassurance and familiarity when shit hits the fan—and boy has it been hitting it lately. Comfort is good. Comfort offers an escape.
It is not wise to always escape.
In this blog I lie to protect my own privacy (it’s even in the FAQ). What I can tell you with certainty is that you can never escape yourself.
I got too bleak and existentialist, so here’s the thing that triggered the whole… thing: seeing a PS2 as part of the 90’s exhibit in A MUSEUM:
Indeed, ‘comfort’ is the operative word here.
ANYWAY. This is an AI newsletter, and I’m digressing a lot. Here’s this week’s content:
In the news: there are some interesting reads around what is causing the AI layoffs, some unfortunate updates, and, to lighten the mood, embodied AI slop.
In the papers: a very solid (and very long) lineup. I particularly enjoyed one about content moderators misclassifying men’s feelings. Even the cut content is good, just… cut because of space. And literally one paper on the lineup that made me think. What a lame week.
News
AI layoffs are looking more and more like corporate fiction that's masking a darker reality | Fortune — ok this is interesting as a read. It does feel (from the article) that it is just an investor-pleasing strategy (‘we are hoping that AI will cover these vacancies’ → hoping being the key bit here).4 There are wrong things with this article, though: basically, if machines replaced humans, ‘output per remaining worker should skyrocket’. They clearly have no idea how overworked we are. But ohkay.
Annual 'Worst in Show' CES list released | AP News — how could this have happened, said companies trying to please investors.
Hundreds of nonconsensual AI images being created by Grok on X, data shows | Grok AI | The Guardian — It’s absolutely terrible. Honestly. But at the same time what did you expect from X/Grok/Elon Musk. There’s zero ethics involved there. My favourite conspiracy theory here is that they do this to stay relevant. They, of course, did not roll it back—they just paywalled it.
Google and AI startup to settle lawsuits alleging chatbots led to teen suicide | AI (artificial intelligence) | The Guardian — remember kids, a settlement is because these absolute twats realised they wouldn’t win. Not because they were doing ‘the right thing’. On the other hand, China is legislating to protect kids from runaway AI. Kinda wish the rest of the planet got their shit together.
Papers
I am not covering anything from the first week of January because anyone with common sense would at least wait until the 7th. Even then I got wayyyy too many good papers.
For funsies I also drew some stats from the first two pages of ArXiv. There are 68 (out of 100) papers using combination titles (Eyecatcher: long sentence); and 47 using product-y names (‘DocDancer’, something-LLM’, ‘ACRONYM’; including the unfortunately named ‘MAGA-Bench’).
Is this why CS has gotten so boring? Every paper sounds the same lol
[2601.02931] Memorization, Emergence, and Explaining Reversal Failures: A Controlled Study of Relational Semantics in LLMs — their terminology is so precise. Exactly what you’d expect from a paper on semantics. Love this.
[2601.04461] Users Mispredict Their Own Preferences for AI Writing Assistance — it’s a good study with interesting findings. Methods-wise, I’m not so sure. It’s great though and I’m just being annoying.
[2601.04730] Automatic Classifiers Underdetect Emotions Expressed by Men — the stat sig values on this one are through the roof (under the basement?). And they cover an exceptionally important subject. You don’t believe me?5 Here’s a good example:
It’s almost as if the whole bottling your feelings because of the pressure of societal expectations could lead to very unhealthy ways to express oneself.
[2601.02671] Extracting books from production language models — lmfao. Aside, we ran the same experiment 3 years ago (and published it, I really should update that arxiv comment). They don’t cite us. Idk what’s with stanford never citing my work.
[2601.04790] Belief in Authority: Impact of Authority in Multi-Agent Evaluation Framework — I do like the merging of semi-soft cheese sciences (psych) with evaluation of LLMs in THIS paper. In general, I find it icky.
[2601.05004] Can Large Language Models Resolve Semantic Discrepancy in Self-Destructive Subcultures? Evidence from Jirai Kei — I really really really really like this paper. Well-motivated, niche, proper investigation, good contribs, etc. That is all.
[2601.00514] The Illusion of Insight in Reasoning Models — this paper is great. The conclusions are fantastic, and also in the related works section their subtle f-you to half of the marketing-driven science (‘LLMs appear to have emergent capabilities’; emphasis theirs).
[2601.01754] Context-Free Recognition with Transformers — this is a very nice paper with great results. The one thing is the assumptions. Turns out there’s a pile of papers showing different results depending on what you assume from the architecture…
[2601.04135] LLMberjack: Guided Trimming of Debate Trees for Multi-Party Conversation Creation — It’s kinda fun, but it really does seem more like a software/framework release than a proper scientific contribution. Science is all about inquiry. That said, there’s some good parametrisation going on, but I wouldn’t consider it a research work.
[2601.01280] Does Memory Need Graphs? A Unified Framework and Empirical Analysis for Long-Term Dialog Memory — you do NOT put on a paper ‘we will release the code soon’. What is this, high school. Christ. Either way, the paper is good and the findings are good and they really should have used the comments section.
[2601.03066] Do LLMs Encode Functional Importance of Reasoning Tokens? — great job, as always. The answer is ‘yes’ (to the title).
In the culture-first department, there is work to add in culture-specific knowledge during training (wouldn’t it be the same as any domain-specific knowledge, I wonder?); and one by my colleagues that is absolutely solid and showing some limits on personalisation (nooooooo…)
There’s also a great systematic eval on 12 African languages. This paper is SO WELL-WRITTEN.6Exemplary work. Also some work in Nawatl dialects and Guarani/Quechua/Aymara. But it is so lazy on the latter that they didn’t even bother removing the ‘code will be released upon submission’. And some work on multilingual LRMs and Tigrinya numbering (if you’re curious).
Finally: [2601.00364] The Role of Mixed-Language Documents for Multilingual Large Language Model Pretraining — ‘translation critically depends on systematic token-level alignments from parallel data’
The Cut Content (due to space)
[2601.04435] Accommodation and Epistemic Vigilance: A Pragmatic Account of Why LLMs Fail to Challenge Harmful Beliefs — you know I love Jurafsky’s work, and I think this is a fun little paper (including the call to pragmatics! It’s so important!). My only concern is that obvious thesis is obvious and resonates heavily of how RLMs work…
[2601.03858] What Does Loss Optimization Actually Teach, If Anything? Knowledge Dynamics in Continual Pre-training of LLMs — ‘these results show that loss optimisation is misaligned with learning progress in CPT [continual pre-training]’
[2601.02043] Simulated Reasoning is Reasoning — this is a POSITION paper (oh, Arxiv). I also disagree with it heavily (but I ran out of space to rant about it). I like it though. It is the only paper this week that made me think.
The Cut Content (I lied)
[2601.01828] Emergent Introspective Awareness in Large Language Models
You know why.
Your circle should be made of people whose comfort films match yours
I remember better the 00s than the 90s, cut me some slack.
I’m even remembering some exec training I attended once, where someone started ranting about the ‘good old days when America was great’ (the 60s). I did interrupt that nitwit to ask about the good old days for whom. Granted that alienating executives is never a good idea, but when you have this authoritative idiot not used to being contradicted just spewing nonsense, all you hope—with all your might—is that you won’t end up like that one day.
I won’t tell you that by inspecting elements you can find the article in outerText.
The number of times I have had to defend this position is dismaying. People should be ashamed of being entrenched in one point of view and not listening to things that challenge their perception of their stupid little world.
I strongly think that you MUST separate results from discussion. I’m not the only one. Literally every other science and most of our subfields do. But I got into some fairly heated arguments during the review of a paper who had the gall to be like ‘the attention is all you need paper didn’t do it’. Like bro because Barry Marshall drank H. pylori it doesn’t make it good science—it was the remaining research and stat sig sample sizes that worked out. They were also just as pleasant about other stuff. Gee, forgive me for being the only one giving you a proper review.







