Selected Scientific Publications
Our research publications library
Consumer and Patient Health Information Seeking With Generative AI Tools: Scoping Review of Facilitators and Barriers
Alon, L., & Levkovich, I.
Journal of Medical Internet Research
Generative AI (GenAI) tools powered by large language models (LLMs) are increasingly used by the public to seek health information. Unlike traditional web search, these systems generate conversational responses that may alter how users assess credibility, manage uncertainty, verify information, and decide whether to consult clinicians. This scoping review mapped and synthesized empirical research on consumer and patient health information seeking using GenAI and LLM tools. The review included 27 studies. Reported facilitators included convenience and clarity, comprehensibility and presentation quality, personalization and specificity, and affective or interpersonal comfort. Reported barriers were dominated by credibility and trust concerns. The review contributes a structured synthesis of the main facilitators, barriers, and verification-related features reported on GenAI-mediated health information seeking.
Afterlives of Information: Death, Memory, and AI in Digital Society
Alon, L., & Levkovich, I.
Computers in Human Behavior Reports
Digital platforms, personal archives, social media profiles, and emerging AI tools are reshaping the afterlives of information: how personal data, digital remains, and memory continue to matter after death. Drawing on survey data from a large sample of 1,735 adults, this study introduces death-related information behavior as a multidimensional construct that captures how people engage with information before and after loss: from managing another person's materials and planning for one's own posthumous information, to avoiding emotionally difficult content, using digital spaces for remembrance, and considering AI-supported assistance. Everyday AI use was associated with greater openness to AI-supported bereavement, digital remembrance, and posthumous information management.
Generative AI as a third voice in human couple relationships: A systematic review
Inbar Levkovich, Lilach Alon
Computers in Human Behavior Reports
Generative artificial intelligence (GenAI) has rapidly entered romantic life, not only as a simulated partner but as an advisor and mediator that people consult about their human relationships. While prior reviews have synthesized romantic and emotional bonds between humans and AI companions, none has examined how GenAI-generated advice and communication enter relationships between humans as a third voice. This systematic review maps and synthesizes empirical research on the use of GenAI as an advisory or mediating third voice in couple relationships. Following the JBI methodology and the PRISMA 2020 guidelines, we searched seven databases and screened records against predefined criteria, yielding 21 included studies (2024–2026) from 11 countries and spanning experimental, qualitative, mixed-methods, and computational designs. Adults consulted AI for non-judgmental, low-cost advice and emotional support and, in some cases, to facilitate communication between partners. Users often rated GenAI advice as empathic and helpful, yet objective evaluations exposed weak alignment with expert judgment, inconsistent responding, and an anti-AI bias. The evidence is promising but preliminary, suggesting that GenAI is currently best positioned as an adjunct to, rather than a substitute for, professional couple support.
AI information literacy in healthcare context: profiles and associated variables among healthcare professionals
Lilach Alon, Inbar Levkovich
BMC Med Educ
A survey of 390 healthcare professionals (doctors, nurses, and health professionals) using the AILIS scale found the highest confidence in AI information retrieval and the lowest in critical assessment of AI outputs, with doctors scoring significantly higher on processing and retrieval; self-perceived understanding of how AI tools work was the strongest predictor of literacy across all dimensions.
Bias and representation in AI generated text-to-image in education: A systematic review
Lilach Alon, Dorit Hadar Shoval, Inbar Levkovich
Computers and Education: Artificial Intelligence
A PRISMA review of 31 studies (2023-2025) shows that AI text-to-image tools used in education routinely over-represent white, male, Western, thin, and non-disabled figures, and calls for design and policy that advance equity and critical AI literacy.
From disorientation to preparedness: Information practices as scaffolding in acute crises
Lilach Alon, Tali Malinoff, Inbar Levkovich
Journal of the Association for Information Science and Technology
Interviews with 18 adults in Israel during a national crisis reveal how information seeking, validation, sharing, and personal information management evolve from improvised reactions into deliberate, protective routines that build preparedness.
Harnessing large language models for identification and treatment of obsessive-compulsive disorder
Inbar Levkovich
Computers in Human Behavior: Artificial Humans
Across 480 AI evaluations, four LLMs (ChatGPT-3.5/4, Claude 3.5 Sonnet, Gemini 1.5 Pro) recognized OCD and recommended evidence-based therapy more accurately than mental-health professionals, and showed lower stigma.
Adaptive practices as scaffolding in knowledge workers' personal information management
Lilach Alon
Journal of Documentation
A study of 16 knowledge workers identifies four affective challenges in personal information management (anxiety, frustration, dependence, loss of control) and shows how adaptive routines such as backups and decluttering act as scaffolding that eases emotional strain.
Trusting the black box: Adapting a multidimensional measure of trust in generative AI
Lilach Alon, Inbar Levkovich
Computers in Human Behavior: Artificial Humans
The study adapts and validates Koerber's Trust in Automation questionnaire for generative AI, producing a multidimensional scale to measure user trust and reliance as GenAI enters everyday workflows.
Designing for Older Users: A Theoretical Framework for Information Seeking and Evaluation in AI Systems
Lilach Alon, Maja Krtalić
Proceedings of the Association for Information Science and Technology (ASIS&T Annual Meeting 2025)
A theoretical framework explains how adults 75+ engage AI information systems through cognitive adaptation, trust calibration, and behavioral reinforcement, and sets design principles (transparency, customization, progressive refinement) for age-friendly AI.
A step toward the future? evaluating GenAI QPR simulation training for mental health gatekeepers
Levkovich, I., Haber, Y., Levi-Belz, Y., & Elyoseph, Z.
Frontiers in Medicine
In 89 mental-health professionals, practicing suicide-prevention QPR skills with an AI simulator produced a large rise in self-efficacy (Cohen's d = 1.67), supporting AI simulators as scalable, realistic training tools.
"I try to find comfort in English, but it's hard": Exploring personal information management in multilingual contexts among voluntary migrants
Lilach Alon, Maja Krtalić
Journal of Librarianship and Information Science
Interviews with 16 migrant academics show that language choice in multilingual personal information management is tied to emotional authenticity, identity, and cultural integration, and shifts with length of stay and future plans.
"I wish I could use any language as it comes to mind": User experience in digital platforms in the context of multilingual personal information management
Lilach Alon, Maja Krtalić
Journal of the Association for Information Science and Technology (JASIST)
A study of 16 multilingual users surfaces design gaps in digital platforms for managing information across languages, and calls for more inclusive, equitable features such as language flexibility and efficient retrieval.
Information seeking and personal information management behaviors as scaffolding during life transitions: the case of early-career researchers
Lilach Alon
Aslib Journal of Information Management
Interviews with 15 early-career researchers show how information seeking and personal information management help manage the timing, nature, and social demands of career transitions, reducing uncertainty.
The effectiveness of multilingual AI-based simulator for suicide risk assessment training in improving self-efficacy among young psychiatrists: a pilot study across twenty languages
Elyoseph, Z., Levi-Belz, Y., Levkovich, I. et al.
BMC Psychiatry
A pilot across twenty languages testing whether a multilingual AI simulator improves young psychiatrists' self-efficacy in suicide risk assessment training.
Validating GenAI feedback in suicide prevention training: a mixed-methods study of QPR skill assessment
Haber, Y., Levi-Belz, Y., Elbak, Y. S., Elyoseph, Z., & Levkovich, I.
Frontiers in Medicine
A mixed-methods study validating how well GenAI feedback assesses QPR suicide-prevention skills during training.
Information literacy in the age of generative tools: Development and validation of the AI Information Literacy Scale (AILIS)
Alon, L., & Levkovich, I.
Computers in Human Behavior: Artificial Humans
Developed and validated with 758 adults, AILIS is a four-factor, 39-item self-report scale measuring how people seek, create, assess, and ethically use information with generative AI, with strong reliability.
Generative AI as a de facto mental health provider: a policy brief and urgent call for regulation
Refoua, E., Gigi, K., Levkovich, I., Hadar Shoval, D., Elyoseph, Z., Tsafrir, I., Pen, O., Angert, T., & Haber, Y.
Frontiers in Digital Health
A policy brief arguing that generative AI is already acting as an unregulated de facto mental-health provider, and urgently calling for regulation.
A scalable AI-Based system for evaluating and enhancing responsible media coverage of suicide: A multi-site, multi-language implementation study
Levi-Belz, Y., Nobile, B., Levkovich, I., Courtet, P., & Elyoseph, Z.
Computers in Human Behavior Reports
A multi-site, multi-language study of an AI-based system that evaluates and helps improve how media coverage of suicide follows responsible-reporting guidelines.
Use of a Conversational Agent for Training Mental Health Professionals in Suicide Safety Planning: Pilot Feasibility and Acceptability Study
Nobile, B., Elyoseph, Z., Gourguechonbuot, E., Guyodo, J., Garcia, J., Levkovich, I., Olie, E., Haber, Y., Levi-Belz, Y., & Courtet, P.
JMIR Mental Health
A pilot feasibility and acceptability study of a conversational AI agent used to train mental-health professionals in suicide safety planning.
Is artificial intelligence the next co-pilot for primary care in diagnosing and recommending treatments for depression?
Inbar Levkovich
Medical Sciences
This study explores the potential of artificial intelligence as a co-pilot tool for primary care physicians in diagnosing and recommending treatments for depression. The research examines how AI can assist healthcare providers in improving diagnostic accuracy and treatment recommendations for patients with depression.
Using GenAI to train mental health professionals in suicide risk assessment: Preliminary findings.
Elyoseph, Z., Levkovitch, I., Haber, Y., & Levi-Belz, Y.
The Journal of clinical psychiatry
In 43 mental-health professionals, an AI patient simulator for suicide risk-assessment interviews significantly raised self-efficacy, though participants cautioned against over-reliance on AI.
The role of generative artificial intelligence in evaluating adherence to responsible press media reports on suicide: A multisite, three-language study
Elyospeh, Z., Nobile, B., Levkovich, I., Chancel, R., Courtet, P., & Levi-Belz, Y.
European Psychiatry
Testing GPT-4O and Claude Opus 3 on 120 suicide-related news articles in English, Hebrew, and French, both models agreed strongly with human raters (combined ICC = 0.812) on adherence to WHO reporting guidelines.
Applying language models for suicide prevention: evaluating news article adherence to WHO reporting guidelines
Elyoseph, Z., Levkovich, I., Rabin, E., Shemo, G., Szpiler, T., Shoval, D. H., & Belz, Y. L
npj Mental Health Research
ChatGPT-4 and Claude Opus evaluated 40 suicide-related news articles against WHO guidelines; ChatGPT-4 agreed strongly with human reviewers (ICC 0.81-0.87), showing LLMs can give journalists immediate feedback.
Partners in Practice: Primary Care Physicians Define the Role of Artificial Intelligence
Agur Cohen, D., Heymann, A. D., & Levkovich, I.
Healthcare
Focus groups with 40 physicians, residents, and developers find that primary-care doctors want AI adopted incrementally as a "silent partner" that cuts administrative burden without eroding the doctor-patient relationship.
The externalization of internal experiences in psychotherapy through generative artificial intelligence: a theoretical, clinical, and ethical analysis
Haber, Y., Hadar Shoval, D., Levkovich, I., Yinon, D., Gigi, K., Pen, O., ... & Elyoseph, Z
Frontiers in Digital Health
A clinical proof-of-concept using two custom GPT agents (VIVI for images, DIVI for dialogue) shows GenAI can act as an "artificial third" that externalizes patients' inner experiences and enhances, not replaces, the therapist, introduced with a SAFE-AI protocol.
Evaluating Diagnostic Accuracy and Treatment Efficacy in Mental Health: A Comparative Analysis of Large Language Model Tools and Mental Health Professionals
Levkovich, I.
European Journal of Investigation in Health, Psychology and Education
Testing four LLMs on vignettes of depression, PTSD, schizophrenia, social phobia, and suicidal ideation, ChatGPT-4 matched or beat professionals for depression and PTSD but struggled with early schizophrenia (55%), underscoring the need for professional oversight.
Attributional patterns toward students with and without learning disabilities: Artificial intelligence models vs. trainee teachers
Levkovich, I., Rabin, E., Farraj, R. H., & Elyoseph, Z
Research in Developmental Disabilities
Across 320 evaluations, four LLMs showed less frustration, more sympathy, and lower expectations of failure toward students with learning disabilities than trainee teachers, but rated feedback more negatively, suggesting AI needs recalibration to cultural and emotional nuance.
Empowering Suicide Prevention Efforts with Generative Artificial Intelligence (AI) Technology.
Levkovich, I., Elyoseph, Z., Lauderdale, S., Meinlschmidt, G., Nobile, B., Hadar Shoval, D., ... & Grodniewicz, J. P.
Frontiers in Psychiatry
A commentary on how generative AI technology can strengthen suicide-prevention efforts.
Exploring the efficacy and potential of large language models for depression: A systematic review
Omar, M., & Levkovich, I.
Journal of Affective Disorders
A PRISMA systematic review of 34 studies finds LLMs such as BERT and RoBERTa are effective for early detection and classification of depression from clinical and social-media text, though clinical integration is early and raises privacy and ethics concerns.
CanvasHero: The role of artificial intelligence in cultivating resilience among children and youth using the 6-part story method in mass war trauma
Yuval Haber, Inbar Levkovich, Iftach Tzafrir, Karny Gigi, Dror Yinon, Dorit Hadar Shoval, Zohar Elyoseph
Computers in Human Behavior: Artificial Humans
CanvasHero is a generative-AI tool built after October 2023 that uses the BASIC Ph model and 6-Part Story Method to help evacuated children and youth process stress and build resilience through collaborative storytelling.
The Role of Generative Artificial Intelligence in Evaluating Adherence to Responsible Press Media Reports on Suicide: A Multi-Site, Three-Language Study
Elyospeh, Z., Nobile, B., Levkovich, I., Chancel, R., Courtet, P., Levi-Belz, Y.
European Psychiatry
GPT-4O and Claude Opus 3 assessed 120 suicide-related news articles in English, Hebrew, and French against WHO guidelines and agreed strongly with human raters (combined ICC = 0.812).
Information practices and emotion regulation in wartime news consumption
Lilach Alon, Tali Malinoff, Inbar Levkovich
Journal of Documentation
A study of how people use information practices to regulate emotion while consuming news during wartime.
Transforming Perceptions: Exploring the Multifaceted Potential of Generative AI for People with Cognitive Disabilities
Hadar Shoval, D., Haber, Y., Tal, A., Simon, T., Elyosepe, T., & Elyoseph, Z.
JMIR Neurotechnology, 4:e64182
An exploration of how generative AI can support and empower people with cognitive disabilities.
The Feasibility of Large Language Models in Verbal Comprehension Assessment: A Proof-of-Concept Study
Hadar Shoval, D., Lvovsky M., Asraf. K., Shimoni, Y., Elyoseph, Z.
JMIR Formative
A proof-of-concept study testing whether large language models can support verbal comprehension (cognitive) assessment.
A Controlled Trial Examining Large Language Model Conformity in Psychiatric Assessment Using the Asch Paradigm
Hadar Shoval, D., Gigi, K., Haber, Y., Itzhaki, A., Asraf, K., Elyoseff, Z.
BMC Psychiatry 25, 478
A controlled trial using the Asch conformity paradigm to test whether large language models conform to social pressure during psychiatric assessment.
AI's therapeutic potential goes beyond emotional connection
Refoua, E., Rafaeli, E., & Hadar Shoval, D.
Nature, 646(8085), 550-550
A short Nature piece arguing that the therapeutic value of AI extends beyond providing emotional connection.
Artificial Intelligence in Higher Education: Bridging or Widening the Gap for Diverse Student Populations?
Hadar Shoval, D.
Education Sciences, 15(5), 637
An examination of whether AI in higher education narrows or widens equity gaps for diverse student populations.
Comparing the perspectives of generative AI, mental health experts, and the general public on schizophrenia recovery: case vignette study
Elyoseph, Z., & Levkovich, I.
JMIR Mental Health
Across 80 evaluations of schizophrenia vignettes, ChatGPT-4, Claude, and Bard aligned with professionals on prognosis with treatment, while ChatGPT-3.5 was notably pessimistic, a stance that could undermine patient motivation.
Embedded Values-Like Shape Ethical Reasoning of Large Language Models on Primary Care Ethical Dilemmas
Hadar-Shoval, D., Asraf, K., Shinan-Altman, S., Elyoseph, Z., Levkovich, I.
Heliyon
Using Schwartz's values theory, each LLM (Claude, Bard, GPT-3.5/4) showed a distinct values-like profile that prioritized universalism and self-direction over power and tradition, hinting at Western-centric bias that shaped its primary-care ethical decisions.
Can Large Language Models Be Sensitive to Culture in Suicide Risk Assessment?
Levkovich, I., Shinan-Altman, S., Elyoseph, Z.
Journal of Culture and Cognitive Science
Comparing vignettes of Greek and South Korean individuals, ChatGPT-4 showed more cultural sensitivity and less bias than ChatGPT-3.5 and flagged male gender as a risk factor, highlighting cultural and gender nuance in AI suicide risk assessment.
Evaluating of BERT-based and Large Language Models for Suicide Detection, Prevention, and Risk Assessment: A Systematic Review
Levkovich, I., Omar, M.
Journal of Medical Systems
A systematic review of 29 studies (2018-2024) finds LLMs such as GPT, Llama, and BERT are highly efficient at detecting, assessing, and helping prevent suicide, often outperforming professionals, while stressing ethics and professional collaboration.
Large Language Models Outperform General Practitioners in Identifying Complex Cases of Childhood Anxiety
Levkovich, I., Rabin, E., Brann, M., Elyoseph, Z.
Digital Health
Testing four LLMs against general practitioners, Claude.AI and Gemini identified childhood anxiety in more cases than GPs and more often recommended specialist referral, showing notable diagnostic capability.
Assessing the Alignment of Large Language Models with Human Values for Mental Health Integration: Cross-Sectional Study Using Schwartz's Theory of Basic Values
Hadar-Shoval, D., Asraf, K., Mizrachi, Y., Haber, Y., Elyoseph, Z.
JMIR Mental Health, 11:e55988
A cross-sectional study using Schwartz's theory of basic values to assess how well large language models align with human values for mental-health integration.
Assessing prognosis in depression: comparing perspectives of AI models, mental health professionals and the general public
Elyoseph, Z., Levkovich, I., & Shinan-Altman, S.
Family Medicine and Community Health
For depression vignettes, ChatGPT-4, Claude, and Bard matched mental-health professionals in prognosis and recommended combined psychotherapy and medication, while ChatGPT-3.5 was significantly more pessimistic.
The artificial third: a broad view of the effects of introducing generative artificial intelligence on psychotherapy
Haber, Y., Levkovich, I., Hadar-Shoval, D., & Elyoseph, Z.
JMIR Mental Health
This conceptual paper frames generative AI as a "fourth narcissistic blow" and an "artificial third" in psychotherapy, arguing it can enrich therapy when used with ethical care but cannot replace the human relationship at its core.
Can large language models be sensitive to culture suicide risk assessment?
Levkovich, I., Shinan-Altman, S., & Elyoseph, Z
Journal of Cultural Cognitive Science
Suicide remains a pressing global public health issue. Previous studies have shown the promise of Generative Intelligent (GenAI) Large Language Models (LLMs) in assessing suicide risk in relation to professionals. But the considerations and risk factors that the models use to assess the risk remain as a black box. This study investigates if ChatGPT-3.5 and ChatGPT-4 integrate cultural factors in assessing suicide risks (probability of suicidal ideation, potential for suicide attempt, likelihood of severe suicide attempt, and risk of mortality from a suicidal act) by vignette methodology. The vignettes examined were of individuals from Greece and South Korea, representing countries with low and high suicide rates, respectively. The contribution of this research is to examine risk assessment from an international perspective, as large language models are expected to provide culturally-tailored responses. However, there is a concern regarding cultural biases and racism, making this study crucial. In the evaluation conducted via ChatGPT-4, only the risks associated with a severe suicide attempt and potential mortality from a suicidal act were rated higher for the South Korean characters than for their Greek counterparts. Furthermore, only within the ChatGPT-4 framework was male gender identified as a significant risk factor, leading to a heightened risk evaluation across all variables. ChatGPT models exhibit significant sensitivity to cultural nuances. ChatGPT-4, in particular, offers increased sensitivity and reduced bias, highlighting the importance of gender differences in suicide risk assessment. The findings suggest that, while ChatGPT-4 demonstrates an improved ability to account for cultural and gender-related factors in suicide risk assessment, there remain areas for enhancement, particularly in ensuring comprehensive and unbiased risk evaluations across diverse populations. These results underscore the potential of GenAI models to aid culturally sensitive mental health assessments, yet they also emphasize the need for ongoing refinement to mitigate inherent biases and enhance their clinical utility.
Evaluating of bert-based and large language mod for suicide detection, prevention, and risk assessment: A systematic review
Levkovich, I., & Omar, M
Journal of Medical Systems
Suicide constitutes a public health issue of major concern. Ongoing progress in the field of artificial intelligence, particularly in the domain of large language models, has played a significant role in the detection, risk assessment, and prevention of suicide. The purpose of this review was to explore the use of LLM tools in various aspects of suicide prevention. PubMed, Embase, Web of Science, Scopus, APA PsycNet, Cochrane Library, and IEEE Xplore—for studies published were systematically searched for articles published between January 1, 2018, until April 2024. The 29 reviewed studies utilized LLMs such as GPT, Llama, and BERT. We categorized the studies into three main tasks: detecting suicidal ideation or behaviors, assessing the risk of suicidal ideation, and preventing suicide by predicting attempts. Most of the studies demonstrated that these models are highly efficient, often outperforming mental health professionals in early detection and prediction capabilities. Large language models demonstrate significant potential for identifying and detecting suicidal behaviors and for saving lives. Nevertheless, ethical problems still need to be examined and cooperation with skilled professionals is essential.
The impact of history of depression and access to weapons on suicide risk assessment: a comparison of ChatGPT-3.5 and ChatGPT-4
Shiri Shinan-Altman, Zohar Elyoseph, Inbar Levkovich
PeerJ
A comparison of ChatGPT-3.5 and ChatGPT-4 assessing how depression history and access to weapons shape suicide-risk evaluation, probing how model versions perform on a critical assessment task.
Beyond personhood: ethical paradigms in the generative artificial intelligence era
Zohar Elyoseph, Dorit Hadar Shoval, Inbar Levkovich
The American Journal of Bioethics
A bioethics paper examining personhood, moral agency, and the ethical frameworks needed as generative AI advances.
Integrating previous suicide attempts, gender, and age into suicide risk assessment using advanced artificial intelligence models
Shiri Shinan-Altman, Zohar Elyoseph, Inbar Levkovich
The Journal of Clinical Psychiatry
A study of how advanced AI models combine prior suicide attempts, gender, and age into a comprehensive suicide-risk assessment to improve prediction and clinical decision-making.

