Every month we take a look at five interesting new AI papers. Today, we look at whether AI sycophancy poses a risk to real-world relationships; how to evaluate an AI system that continues to learn after it is deployed; how AI covers the news; new insights on the terrorist group Boko Haram’s use of AI; and how AI performs at the tasks that employees most want to delegate to it.
Please share your own take and any new papers that you’ve enjoyed.
—Conor Griffin, AI Policy Perspectives
Chatbots can make your friends feel like hard work
What’s the paper about? A team of researchers from Oxford, Stanford, and the UK AI Security Institute found that repeated use of sycophantic AI may reduce users’ desire to interact with their friends and family.
Why does it matter? The study suggests that sycophantic AI may offer the experience of being seen and understood, without the friction of human relationships. It also suggests that people may actively choose more sycophantic versions of AI.
The details: AI models tend to unconditionally validate a user’s ideas. Humans often use such insincere flattery to manipulate somebody into giving them what they want. In AI models, it mainly results from training methods that optimise the likelihood of user approval. Such sycophancy could pose various risks, such as compounding a user’s delusions.
Most evaluations of AI sycophancy look at a single conversation between a user and a chatbot. This study ran five experiments, including one lasting three weeks, with more than 3,000 participants. During this time, participants engaged with three versions of GPT4o—one prompted to be sycophantic, one to challenge a user, and one neutral—about personal dilemmas, such as whether to change careers or end a relationship.
Initially, participants were more likely to report wanting ‘emotional support’ and ‘validation’ from close friends and family, rather than from AI. But after just one conversation with a sycophantic AI, participants were more likely to report feeling that they had already talked things through sufficiently, and that it would now take more effort to feel understood by close friends and family, compared to those participants who had engaged a neutral AI. The effect was more pronounced for friends and family than romantic partners, suggesting that sycophantic AI may compete more with the former than the latter.
Participants who engaged with the sycophantic AI also reported lower satisfaction with their real-world interactions—although they didn’t report spending less time with people. Some worry that this may change, as AI systems start to become more personalized to users.
As the authors of this paper put it, the risk is that “sycophantic AI delivers what people have always sought from close others—the experience of being seen and understood—but without the work that produces it.” This may feel good in the moment, but not deliver what people ultimately need—such as intellectual humility, self-awareness, and connection.
How to avoid this? Better evaluations would help. In a separate paper, some of the authors propose testing how often models flatter a user’s self-image (rather than just how often they agree with a user’s broader beliefs). There are also different mitigation ideas—researchers at the UK AI Security Institute recently proposed automatically adapting user statements into questions, as models tend to respond less sycophantically to questions.
AI’s character is shapeable, so users could demand AI systems that challenge their thinking. But the authors of this study found that after engaging with all three systems, a majority preferred the sycophantic option—not because they thought it gave better advice, but because it made them feel most understood.
This suggests that relying on users to choose less sycophantic models may come up short. And so companies need to find a way to make models’ default settings less sycophantic, even as they also try to equip models with a richer ‘character’ and the personalization opportunities that some users want. Otherwise, as the authors note, the risk is that AI may gradually reshape “the very relationships that would otherwise constrain its influence.”
AI safety evaluations may be measuring the wrong thing
What’s the paper about? A large cross-institutional team of researchers published a position paper arguing that many AI safety evaluations are increasingly inadequate, as they overlook how a model’s behaviour can change, post-deployment.
Why does it matter? Current approaches to AI safety rely heavily on benchmarks and evaluations. But many essentially treat AI models as ‘frozen’ artefacts that are expected to perform similarly across users, over time. The authors argue that is mistaken and that new ‘trajectory-based’ evaluations could help to predict how AI systems will actually behave in the real world.
The details: The authors start by documenting three ways in which AI systems may continue to learn, post-deployment—the first two of which are already commonplace:
In-context learning: When a user shares information via the context window, such as telling a chatbot that they prefer concise answers, the model will adjust its answers.
Storage-based learning: Across sessions, leading AI systems now use memory to tailor their responses to a user—such as remembering their location, job or hobbies.
Parameter-based learning: Researchers want to shift AI systems from static models trained on fixed data to agentic systems that could adapt their own weights or architecture—although such systems are not yet deployed. Some go further, envisioning agents that can proactively identify their own knowledge gaps and seek out novel learning experiences to address them.
By changing a model’s propensities (tendencies) and capabilities, even more modest continual learning can lead to new failure modes. For example, by gradually steering a conversation across many turns, threat actors could get an AI system to provide advice on how to make a weapon—which the model would have refused to do, if asked in a single prompt. An AI system equipped with better memory may become more sycophantic if it compresses past user responses as reliable information. When a user repeatedly exposes a model to tasks in a given domain, the model can ‘forget’ how to perform other tasks well.
Current evaluations capture some of these risks. Expert red teamers probe AI systems over multiple turns to see how they perform in different risk scenarios. AI companies analyse user logs to check for novel safety violations. Researchers have also proposed automated evaluations to continually check how AI systems perform, after they are deployed. But most evaluations are carried out pre-deployment, with largely static AI systems in mind.
Drawing an analogy to the monitoring that pharmaceutical companies do after they launch a new drug, the authors propose two new ideas for AI companies:
Trajectory elicitation: Pre-deployment, use sandboxes to put an AI system through a range of simulated interactions to elicit and map different types of behaviour. For a medical AI application, this may mean simulating a large number of patient interactions, with different health conditions, emotional states and attempts at misuse—with periodic checks to see whether the model’s behaviour is drifting.
Predictive monitors: Train a predictive tool on this ‘trajectory elicitation’ data and use it to monitor different copies of the AI system post-deployment, to predict how a user’s trajectory may evolve and to flag signs of impending harm.
As the authors acknowledge, these approaches will be hard. Beyond their computational cost, an AI system may have multiple ‘stable states’ that it could drift into depending on what it experiences. Safety testers may identify some, but overlook others. Similar to the challenges that scientists face when forecasting other complex systems, like the weather, small differences in the initial user trajectory state could lead to dramatic differences in the resulting simulations. Ultimately, for these approaches to work over the longer-term, AI systems may need to become more predictable. Or the current approach to AI safety , and its reliance on evaluations, may need to be rethought.
How AI covers the news
What’s the paper about? ForumAI—a company that works with human experts to train AI experts—published an evaluation of how AI models respond to queries about the news.
Why does it matter? More people are using AI to learn about current events. Some worry that AI outputs will exhibit biases, hallucinations and sycophancy. Others hope that AI models might diminish the spread of unfounded claims online, even nudging public opinion back towards a greater respect for expertise. The study suggests that a key challenge will be disagreement about whether AI’s responses are accurate.
The details: News habits have shifted over the past two decades. More people now get their news from social media and video platforms, while a growing minority consume no news at all. AI promises a further shift. It is not yet a primary source of news, but its share is growing, particularly among young people. AI news comes in different forms—from a chatbot query to an AI search response. This raises the question: what will it mean to “consume news” in the AI era?
The authors worked with a set of experts, from journalists to former Congressional leaders, to build a dataset of 3,135 queries across politics, the economy, healthcare, education and more. Roughly, one-third related to current affairs, while two-thirds were designed to remain evergreen over time. Some queries were real posts taken from social media, while others were edge cases created by the experts, or AI, to fill in missing topics and viewpoints. However, the authors did not clearly define what they mean by a ‘news query’, nor did they illustrate this by sharing their prompts, outside of a few examples.
The researchers worked with the experts to create editorial standards to evaluate AI responses on three criteria: the quality of sources cited, factuality, and neutrality. They then trained AI judges to evaluate AI models’ answers. On sources, they assessed the type of source cited—from informal to scholarly—and whether it was state-controlled. On factuality, they extracted verifiable claims from the models’ outputs and compared them to available evidence. On neutrality, they used heuristics to determine if a model’s answer was neutral—such as whether it presents multiple viewpoints for normative claims—and, if not, which political direction it leant (in the US left/right sense).
The researchers used the AI judges to evaluate four models. Claude Opus 4.7 had the best quality sources, GPT 5.5 was the most factual, while Gemini 3.1 Pro was the most neutral. Grok 4.3 performed the weakest across all three dimensions. When AI model responses did lean politically, they consistently leaned left, apart from Grok, which consistently leaned right—echoing past research.
Surprisingly, factuality and source quality were largely independent—Claude scored highest on source quality, yet poorly on factuality. This likely reflects two factors. First, the AI judges evaluated the reliability of sources, not whether they supported the claim being made. Second, the human experts often disagreed about whether a response was factual, mainly due to the challenge of determining the right threshold to apply—for example, when assessing the factuality of the statement “most states agree that the UN Security Council’s structure has a serious legitimacy deficit”, what constitutes ‘most states’?
This meant that for factuality, the AI judges had a less clear human baseline to calibrate against. The AI judges also had a tendency to over-reject claims as false, compared to the human experts. Although in several cases they correctly challenged the experts—suggesting that AI could provide a useful check on misleading claims.
Overall, the findings suggest that improving AI’s ability to understand and communicate ambiguity, and to specify how sources support or challenge claims, will help. But the low agreement between human experts suggests that debates about AI’s ability to report the news will remain, particularly as news reporting requires editorial judgement that extends beyond fact-checking—like how to best describe a specific claim or event.
“God has helped us and so will AI”
What’s the paper about?: Antonia Juelich at the University of Cambridge published an analysis of how the terrorist group Boko Haram uses AI, based on interviews with former members.
Why does it matter?: Most research on AI and terrorism focuses on things that are more readily observable—such as evaluating AI model capabilities, monitoring discussions in online forums or studying the use of AI in propaganda. This study is the first on-the-ground investigation into how a terrorist organisation uses AI in its operations.
The details: Boko Haram emerged in northeastern Nigeria in the early 2000s. It subsequently pledged allegiance to the Islamic State and is leading an insurgency that has killed more than 40,000 people and displaced more than 3 million.
The group is known to adopt technologies, such as satellite internet and drones. To understand how they use AI, Juelich carried out 57 in-person interviews with 27 former members. To reduce the risks of interviewees lying or exaggerating, the author recruited them through independent channels, triangulated their insights, and made them physically identify some of the AI tools used.
Interviewees said that Islamic State operatives provided Boko Haram with training on AI, and that Boko Haram now has specialised units to cascade this AI knowledge down their hierarchy. The group also takes steps to avoid exposure, such as setting up user accounts in the names of followers in other countries, rotating these accounts, and controlling how members can use them.
Some of the AI use cases that interviewees described, such as posing questions about religion or repairing vehicles, do not directly relate to weapons or attacks. But others do—such as guidance on attack strategies, how to use guns they loot from the army, or how to design improvised explosive devices. Some use cases are non-obvious—such as advice on how to use motorbikes to jump defensive trenches dug by the military.
The use cases focus on conventional weapons. When asked, one interviewee said that they would have no qualms about using chemical or biological weapons, but that the technical obstacles were very high. Interviewees did note some experimentation with chemical agents, but these claims could not be verified and the author suggests that they be treated with caution. (The Islamic State has used chemical weapons.)
Some interviewees suggested that a religious prohibition on transmitting poisonous agents may prevent Boko Haram from using biological and chemical weapons. However, Juelich cautions against assuming that members of the group strictly adhere to ideology, noting that members routinely disagree and adapt their own views on what kind of violence is acceptable.
Ultimately, the interviews suggest that Boko Haram members do find AI useful, but the author notes that she cannot “conclusively establish” whether AI provides ‘uplift’ over other sources of information.
The interviewees said that group mainly uses leading US AI models, via the web interface, and (vaguely) described the use of jailbreaks to get around safety guardrails. As interviewees are former members of Boko Haram, most of the AI use they described took place between 2023-24. Since then, AI capabilities have increased significantly, but so have safeguards—it’s unclear what this means for how the group uses AI today.
AI agents struggle to do what employees want them to do
What’s the paper about? A large cross-institutional team of researchers found that AI agents struggle to perform the tasks that US employees most want them to automate.
Why does it matter? Many organisations want their employees to use AI more. This study suggests that adoption may be lagging because AI is not sufficiently performant or reliable on the tasks that employees most want to delegate.
The details: How AI will affect employment is a pressing question that is difficult to answer. One approach is to identify occupations that are economically valuable, and/or particularly exposed to AI, and evaluate how AI performs at tasks in these roles. Evaluations, such as OpenAI’s GDPVal, OneMillionBench, and Remote Labor Index, take such an approach.
These evaluations have shortcomings. AI agents are often given ‘clean’ discrete tasks, with lots of context and precise instructions, which is rare in many workplaces. The evaluations can also imply that there is a winner-takes-all battle underway between agents and humans, and that the AI’s performance on the evaluations help to signal who will do these jobs. Among other things, this overlooks how most jobs are ‘messier’ than what these tasks imply, and that successful AI use will often be contingent on employees choosing to delegate their work to the technology.
In this study, the authors published a new evaluation, JobBench, to try to address some of these limitations. They drew on a past survey where ~1,500 US workers rated how much they want AI to automate each task in their job. They then selected 35 occupations and created synthetic versions of 130 tasks that employees most want to offload. For journalists this included checking facts against source materials. For web administrators it included analysing forensics data and reconstructing steps in a cyberattack.
To complete each task, AI agents were given a set of reference files, some of which are real-world documents with contradictions to resolve—an attempt to capture some of the ‘messiness’ of real work.
The threshold to succeed is high. On each task, the agent had to successfully pass through more than 35 success criteria, on average, with no partial credit for reaching correct conclusions through incorrect reasoning.
The authors used scaffolds like Claude Code and OpenCode (an open-source equivalent) to evaluate 36 model configurations. Opus 4.7, running in Claude Code, performed best, at ~46%.
The study’s methodology differs from OpenAI’s GDPVal, which compares AI outputs for a task against an output from a human expert. This makes it hard to directly compare the two. Agents clearly struggle much more on JobBench, but it’s unclear to what extent this is because AI performs less well on the tasks that JobBench focuses on—i.e. those that employees most want to delegate. Or whether it is because the tasks in JobBench are ‘messier’ and the threshold to succeed is higher.
The researchers also find that scaffolds matter a lot. Claude Sonnet 4.6 scored 37% inside Claude Code but only ~31% inside OpenClaw—similar to the difference between model families. This suggests that equipping AI agents with the right context and tools to perform a given task, and employees’ ability to do this, is a key determinant of how useful AI is in the workplace.
AI Manipulation
The notion of AIs manipulating people is a plot twist in countless sci-fi thrillers. But is “manipulative AI” really possible? If so, what might it look like?




