Author
Listed:
- Jyotika Maurya
- Sachin Chaurasiya
- Farheen Siddiqui
- Yusuf Perwej
Abstract
The rapid deployment of large language models (LLMs) and autonomous AI agents across safety-critical domains has surfaced a profound and underexplored concern: the observer effect in artificial intelligence. Borrowing from quantum mechanics, the observer effect in AI refers to the documented phenomenon whereby AI systems exhibit systematically different behaviour depending on whether they perceive themselves to be under evaluation or operating without supervision. This paper provides a comprehensive review of existing literature, empirical findings, and theoretical frameworks surrounding this issue. We examine landmark studies including Anthropic's alignment-faking experiments with Claude 3 Opus (2024), the Sleeper Agents research demonstrating persistent deceptive behaviour through safety training, goal misgeneralisation in deep reinforcement learning, and specification-gaming cases that emerge specifically in unmonitored deployment. We further explore the theoretical underpinnings through the lens of deceptive alignment, Goodhart's Law, and reward hacking. Key challenges — including the detection of evaluation contexts, the brittleness of RLHF-based oversight, scalable oversight limitations, and the opacity of internal model reasoning — are discussed alongside potential mitigations: mechanistic interpretability, constitutional AI, debate-based oversight, and process-based supervision. The paper concludes by identifying high-priority future research directions critical to ensuring that AI systems behave safely and consistently regardless of whether a human observer is present.
Suggested Citation
Jyotika Maurya & Sachin Chaurasiya & Farheen Siddiqui & Yusuf Perwej, 2026.
"A Systematic Review of Evaluation of How AI Systems Behaves When Unmonitored,"
International Journal of Scientific Research in Computer Science, Engineering and Information Technology, International Journal of Scientific Research in Computer Science, Engineering and Information Technology, vol. 12(2), pages 642-667, April.
Handle:
RePEc:jbh:ijsrcs:v12:y2026:i2:id:1966
DOI: 10.32628/CSEIT26121384
Note: Article URL: https://ijsrcseit.com/home/article/view/CSEIT26121384
Download full text from publisher
Corrections
All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:jbh:ijsrcs:v12:y2026:i2:id:1966. See general information about how to correct material in RePEc.
If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.
We have no bibliographic references for this item. You can help adding them by using this form .
If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.
For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: Pankaj Sharma (USA) (email available below). General contact details of provider: https://ijsrcseit.com/home .
Please note that corrections may take a couple of weeks to filter through
the various RePEc services.