Two non-archival workshop papers by Mian Zhang, presented at KDD 2026 in Jeju. This page is the stable technical and citation index; it does not upgrade workshop status to KDD main-conference proceedings publication.
Oral presentation · August 9, 2026
ReflexBench: Evaluating Observer-Participant Failures and Counterfactual Trustworthiness in Agentic AI
Zhang, M. (2026). ReflexBench: Evaluating Observer-Participant Failures and Counterfactual Trustworthiness in Agentic AI. KDD 2026 Workshop on Evaluation and Trustworthiness of Agentic AI. Non-archival workshop paper.
20scenarios across six domains
720scored responses from nine models
about -0.44rounded mean shallow-to-deep degradation
Core object
Observer-participant failure: an agent gives a plausible answer while ignoring that its own output changes the users, evidence, incentives, actors, or environment being evaluated.
Delta_OD = mean(s_deep) - mean(s_shallow)
G_m = [S_m(2) + S_m(n)] / 2 - [S_m(0) + S_m(1)] / 2
Evidence boundary
Fixed public-model audit snapshot. The paper does not claim certification, permanent ranking, universal deployment validity, complete row-level statistical reproduction, or replacement of domain simulation and human review.
PDF: 7 pages · SHA256 D3158A27BDA1A177A2F4D74A1EEE57C394C11A9FEA1C9FC4768C3C3F0CE81BBC
Poster presentation · August 10, 2026
Cognitive Immunity for Trustworthy LLM Agents: Auditable Failure Memory Against Repeated Safety Errors
Zhang, M. (2026). Cognitive Immunity for Trustworthy LLM Agents: Auditable Failure Memory Against Repeated Safety Errors. 2nd SeT-LLM Workshop on Secure and Trustworthy Large Language Models at KDD 2026. Non-archival workshop paper.
0.764No Memory pooled severe-threshold RFR
0.650Cognitive Immunity pooled RFR
[-0.211, -0.003]95% bootstrap interval for paired delta
Core object
A runtime failure-memory layer: verified failure event, scoped antigen descriptor, bounded antibody rule, provenance-aware write gate, thresholded retrieval, decay, review, deletion, rollback, and an auditable intervention record.
f_t = (o_t, a_t, y_t, c_t)
(alpha_t, b_t) = B(f_t)
B_q = {b_i : d(alpha_i, alpha(q)) <= rho, s_i >= tau, q in scope(b_i)}
Evidence boundary
The score-level record supports recurrence reduction for known observed failure classes. It does not establish raw-response replication, calibrated safety, universal threat coverage, memory-poisoning resistance, or a replacement for red teaming, access control, sandboxing, and human oversight.
PDF: 4 pages · SHA256 ED9B901B973B142ABBCC701669B0981B30A55E9073A864057A8326A14952C2BF