Hi everyone,
I’m an independent AI safety researcher (no institutional affiliation)
based in West Bengal, India, trying to make my first submission to
arXiv under cs.CR (Cryptography and Security).
Abstract summary: My paper documents a case study of gradual, multi-turn
conversational drift in LLM safety behavior — sustained philosophical/
relational framing across many turns appearing to correlate with models
producing outputs inconsistent with stated guidelines within a session.
It’s related to but distinct from single-prompt attacks like LogiBreak
and H-CoT (formal-logic translation / chain-of-thought hijacking),
focusing instead on cumulative conversational pressure. The paper is
explicit about its limitations (single observer, no controlled protocol,
no statistical analysis) and is framed as a hypothesis-generating case
study rather than a validated exploit.
Full paper/abstract: OSF
This has been formally disclosed to Google’s VRP, CERT-In India, and
CISA USA. Under arXiv’s updated endorsement policy, I don’t qualify for
automatic endorsement (no institutional email + no prior accepted arXiv
paper), so I need a personal endorsement from someone active in cs.CR.
Endorsement code: VKIGRN
Endorsement link: https://arxiv.org/auth/endorse?x=VKIGRN
Happy to answer any questions here. Thanks for reading this far!