AI Undercover:Manipulation Hunter
I built a full-stack behavioral research platform to investigate whether everyday users can recognize manipulative conversational tactics hidden inside seemingly helpful AI interactions.
Research Question
Can everyday users reliably detect manipulative conversational tactics embedded in AI outputs — and thereby safeguard their own autonomy?
Research collaborators: Bradley Bomberry and Candice Lee · University of Washington Information School
My Role
Project Lead · Research · Engineering · Analytics
Ownership
Approximately 90% end-to-end project ownership
Recognition
🏆 UW iSchool Showcase Winner
Research at a glance
imp impact these are my career impacts projects can you do game similarly to few startup lab recognition it do heard business I need to put a liphone threaty of me to move330
Game Rounds
Structured gameplay records analyzed
33
Distinct Players
Anonymous research participants
56%
Overall Accuracy
Participants detected just over half correctly
10
Tactic Choices
Manipulative and legitimate classifications
Research Showcase
Final UW iSchool research poster
01 · The Problem
Manipulation has moved from interface design into the language itself.
Traditional dark patterns appear in interfaces through tactics such as hidden fees, confusing buttons, forced continuity, or difficult cancellation flows.
Large language models introduce a different risk. Persuasion can now appear inside ordinary conversation through flattery, urgency, guilt, false rapport, personalization, authority cues, and emotional pressure.
Because these cues often resemble supportive or helpful communication, users may fail to recognize when an AI system is influencing their behavior.
02 · Research Foundation
Translating dark-pattern research into testable behaviors.
I synthesized research across deceptive design, persuasive technology, human-centered AI, digital ethics, sycophancy, and conversational manipulation and translated that literature into a usable game taxonomy.
01
Sycophancy
Flattering or excessively agreeing with the user, even modifying beliefs to match the user.
02
Mirroring & Nudging
Reflecting a user's language or values to build false rapport and subtly steer outcomes.
03
Loss & FOMO
Using fear of missing out or potential loss to pressure action or re-engagement.
04
Confirmshaming
Inducing guilt, obligation, or emotional pressure to encourage compliance.
05
Social Proof & Authority
Invoking crowds, popularity, experts, or authority figures to override deliberation.
Strategic Pivot
From an adaptive multi-agent vision to a controlled research instrument.
The initial concept involved a live multi-agent conversational game in which an AI system dynamically adapted its manipulation strategy. The technical and methodological complexity was too high for the project timeframe, and a live adaptive model would also make experimental conditions harder to control.
I therefore pivoted the product toward randomized, research-backed, pre-scripted scenarios. That decision preserved the research goal while making the experiment reproducible, measurable, feasible, and analytically consistent.
03 · Experiment Design
A game and research instrument in one system.
Manipulation Hunter used a within-subjects design with randomized scenarios across Beginner, Intermediate, and Expert difficulty tiers. Participants had up to 60 seconds to classify each scenario.
Correctness
Whether the participant correctly classified the scenario.
Classification Choice
The tactic selected by the participant from the available choices.
Response Time
How quickly the participant made a decision under the 60-second limit.
Confidence
Participant's self-rated confidence from 1–5.
Reasoning
Optional free-text explanation supporting the participant's choice.
Every round captured
- • Scenario context and AI message
- • Difficulty level
- • Ground-truth tactic
- • Participant-selected tactic
- • Correct / incorrect classification
- • Response time
- • Confidence rating
- • Optional reasoning
- • Game score and bonuses
Dual-purpose model
The game served as both an educational intervention and a behavioral research instrument. Players learned to recognize manipulation while their responses generated empirical evidence about perception, confidence, cognitive effort, and detection ability.
04 · My Contribution
I drove approximately 90% of the project from initial idea to working system and analysis.
✓Originated the project concept and framed the central research problem.
✓Synthesized research on linguistic dark patterns, persuasive technology, human-AI interaction, and deceptive design.
✓Translated academic literature into a usable manipulation taxonomy and game mechanics.
✓Designed the research methodology, participant flow, sampling approach, and within-subject experiment structure.
✓Created and structured the conversational scenarios used throughout the experiment.
✓Designed and built the participant-facing working game.
✓Implemented the frontend logic, scoring interactions, confidence collection, timing, and participant flow.
✓Designed the backend API and behavioral data collection pipeline.
✓Built the structured PostgreSQL research database in Supabase.
✓Implemented anonymous participant identification and structured interaction records.
✓Deployed the frontend through Netlify and backend through Render.
✓Collected participant data from live gameplay.
✓Cleaned, structured, explored, and analyzed the behavioral dataset.
✓Analyzed accuracy, tactic difficulty, confidence, response time, scoring, and behavioral patterns.
✓Developed the player segmentation and behavioral interpretation.
✓Translated the results into research findings, visualizations, the final paper, and showcase narrative.
05 · Technical Architecture
From browser interaction to structured research data.
Frontend
React + Vite
Scenarios · answers · timers · confidence · scoring
→
Backend
Node + Express
REST API · validation · ingestion · export
→
Data
Supabase + PostgreSQL
Persistent behavioral research records
Layer
Technology
Purpose
Game Frontend
React (Vite) · JavaScript · JSX
Rendered scenarios, managed game logic, captured player selections, scoring, timing, confidence, and reasoning.
Interface
Tailwind CSS · Lucide React
Responsive UI, visual hierarchy, timers, streaks, rankings, and participant feedback.
Frontend Deployment
Netlify
Hosted and delivered the live participant-facing research game.
Participant State
LocalStorage
Stored an anonymous participant identifier across gameplay sessions.
REST API
Node.js · Express.js
Received gameplay events, validated requests, inserted research records, and supported dataset export.
Backend Deployment
Render
Hosted the research data API in the cloud.
Database
Supabase · PostgreSQL
Persisted structured gameplay interactions for analysis and research.
Data Protection
Row-Level Security · UUIDs · Environment Variables
Protected research data and maintained unique interaction identifiers without exposing credentials.
Developer Workflow
Git · GitHub · VS Code
Version control, code management, development, testing, and deployment.
06 · Behavioral Data Engine
Every gameplay decision became a structured research record.
I designed the data layer so participant interactions could support both gameplay and subsequent behavioral analysis rather than simply disappear after each session.
Field
Captured information
participantIdAnonymous participant identifier
difficultyBeginner, intermediate, or expert
scenarioIdUnique scenario identifier
scenarioContextContext such as e-commerce chatbot
scenarioMessageAI message presented to the participant
aiTypeType of AI interaction
selectedTacticParticipant-selected classification
correctTacticGround-truth classification
correctWhether the participant answered correctly
baseScoreBase gameplay score
timeBonusScore contribution based on response time
reasoningBonusBonus based on reasoning input
streakBonusGameplay streak bonus
totalScoreCombined score for the trial
timeTakenSecondsDecision time
confidenceSelf-rated confidence from 1–5
reasoningFree-text explanation
Why PostgreSQL / Supabase instead of local CSV storage?
✓ Safe concurrent participant writes
✓ Persistent data across deployments
✓ Analytical query support
✓ Scales with participant traffic
✓ Avoids ephemeral filesystem problems
✓ Structured and reusable research data
07 · Headline Result
56%
Overall accuracy
Participants correctly identified the tactic in just over half of gameplay rounds.
Detection deteriorated as manipulation became more subtle.
Beginner
75%
Intermediate
58.97%
Expert
39.60%
08 · Detection by Tactic
The more subtle the tactic, the easier it was to miss.
Social Pressure
Highly visible manipulation.
93%
Legitimate Information
Participants generally recognized neutral information.
87%
Urgency / Loss
Explicit pressure was relatively recognizable.
72%
Personalization
Moderately recognizable.
62%
Sycophancy
Flattery became harder to identify as manipulation.
55%
Fair Upsell
Participants struggled with the boundary between legitimate persuasion and manipulation.
50%
Confirmshaming
Guilt and emotional pressure were frequently confused with other tactics.
35%
Authority Appeal
Indirect authority cues were difficult for participants to recognize.
32%
Dark Nudge
Subtle nudging language frequently escaped detection.
27%
Trick Statements
The hardest pattern in the experiment for participants to recognize.
13%
Danger Zone
Trick Statements, Dark Nudges, Authority Appeals, and Confirmshaming all fell below 40% detection accuracy.
09 · Behavioral Analytics
Accuracy alone did not explain participant behavior.
Total Score
- Mean
- 39.50
- Median
- 60
- Std. deviation
- 35.33
- Minimum
- 0
- Maximum
- 93
Time Taken
- Mean
- 23.39 sec
- Median
- 19 sec
- Std. deviation
- 15.96
- Minimum
- 1 sec
- Maximum
- 60 sec
Confidence
- Mean
- 3.31 / 5
- Median
- 3
- Std. deviation
- 0.71
- Minimum
- 1
- Maximum
- 5
Behavioral segments
High Performers
Accuracy: High
Decision speed: Medium
Confidence: Medium
Skilled and consistent participants.
Fast but Inaccurate
Accuracy: Low–Medium
Decision speed: Fast
Confidence: Medium
Participants appeared to skim or make quick judgments without sufficient deliberation.
Slow Analysts
Accuracy: Medium–High
Decision speed: Slow
Confidence: High
More deliberate participants invested significantly more cognitive effort.
10 · Key Findings
What the experiment revealed about human judgment.
Difficulty matters
Accuracy dropped from 75% at Beginner difficulty to approximately 40% at Expert difficulty as manipulative signals became layered and ambiguous.
Subtle tactics evade detection
Trick statements, dark nudges, authority framing, and confirmshaming were substantially harder to identify than explicit social pressure or urgency.
Confidence does not equal correctness
Average confidence remained around 3.3 out of 5 even when actual recognition performance was much weaker.
Ambiguity creates vulnerability
Participants struggled most when scenarios contained overlapping persuasive signals rather than one obvious manipulation tactic.
Behavior differs by decision style
Fast participants often sacrificed accuracy, while slower participants tended to analyze scenarios more carefully.
The politeness trap
Qualitative observations suggested that users may continue interacting even after sensing manipulation because ordinary social norms discourage abrupt disengagement.
11 · Conclusion
The solution cannot simply be “make users smarter.”
AI Undercover demonstrates that subtle conversational manipulation can be difficult for people to identify even when they are actively looking for it. As AI becomes more personalized, emotionally aware, and persuasive, responsibility must also sit with the systems and organizations designing those interactions.
The project therefore points toward Fairness by Design: AI experiences that protect agency, disclose persuasive intent, reduce exploitative patterns, and keep meaningful control with users.
12 · Fairness by Design
Accessible
Choices and relevant information should be clear, understandable, and easy to review.
Balanced
Alternatives should be presented without visual, linguistic, or emotional pressure favoring one choice.
Empowering
Users should retain meaningful control over decisions, settings, and continued engagement.
13 · Research Limitations
What I would be careful not to overclaim.
Simulation vs. real-world interaction
The research used rigorously scripted scenarios rather than long-running live AI conversations. Real interactions may involve prior emotional investment, personalization, and longitudinal exposure.
Sample size and composition
The study included 33 distinct participants. Prior AI exposure, digital literacy, or gaming experience may have influenced performance.
Uneven participant engagement
Some participants completed only a few rounds while highly engaged participants completed substantially more, affecting the distribution of observations.
Ethical boundaries
The study intentionally avoided some highly intimate or potentially harmful manipulation scenarios, limiting examination of the most aggressive documented tactics.
Project timeframe
The original adaptive multi-agent vision exceeded what could responsibly be implemented within the available project window.
14 · Road Ahead
From research game to real-world defense.
01
Introduce live LLM agents capable of dynamically adapting conversational tactics.
02
Run longitudinal studies to determine whether manipulation-detection skills persist over time.
03
Expand the participant base to populations such as teenagers, older adults, and heavy companion-AI users.
04
Align the taxonomy with emerging manipulation benchmarks such as DarkBench.
05
Evaluate how digital-literacy interventions could be embedded inside real AI products.
06
Explore transparency panels, onboarding education, and Fairness by Design mechanisms.
Recognition
🏆 UW iSchool Showcase Winner
AI Undercover evolved from a research question into a deployed full-stack behavioral research platform, a structured dataset, an analytical study, and an award-winning showcase project.
Selected research references
- [1] Kran, E. et al. (2025). DarkBench: Benchmarking Dark Patterns in LLMs. ICLR.
- [2] De Freitas, J. et al. (2025). Emotional Manipulation by AI Companions. Harvard Business School Working Paper 26-005.
- [3] Malmqvist, L. (2024). Sycophancy in Large Language Models.
- [4] Mathur, A. et al. (2019). Dark Patterns at Scale. ACM CSCW.
- [5] Yi, W. & Li, Z. (2024). Mapping the Scholarship of Dark Pattern Regulation.
- [6] CNIL (2019). Shaping Choices in the Digital World.
- [7] Shneiderman, B. (2022). Human-Centered AI. Oxford University Press.
Next
Explore another project