Chat- IRB? How Application-specific Language Models can Enhance Research Ethics Review
An interview with Joel Seah
DISCLAIMER
Any questions you may have about legal information within this article should be discussed with your attorney or legal counsel at your institution. SPIRIT is an educational resource ONLY and is not meant to substitute research ethics oversight or approval.
The Paper that Took LinkedIn Through the Roof
I first saw Joel Seah on LinkedIn proudly announce his first authorship paper titled “Chat-IRB for LMICs: an opportunity for ethics review capacity-building and protection against ethics dumping, IRB shopping, and other exploitative research practices—a response to Moodley et al”.
Here, we discuss his perspective on this collaborative effort on a deeper level.
Your paper illustrates how AI can be used as a tool for research oversight. Where exactly do you think the boundary lies between AI assistance and human moral participation?
Depending on how the IRB-user intends to deploy it, the utilisation of “Chat-IRB” — application-specific large language models for research ethics review — can vary considerably. Its uses can range from relatively simple or “mundane” applications to more substantive ones.
For example, simple or ‘low-hanging fruit’ applications like: assisting IRB staff (and/or members) in pre-review screenings, and autonomously generating summaries of IRB applications/study protocols; to more substantive functions like: prompting the IRB to reconsider or provide additional justifications for proposed decisions that deviate from past precedents or ethical guidelines, and providing decisional support to inform the IRB’s ethical analysis on specific issues. And the latter would invariably involve AI outputs that can influence the IRB’s ethical judgements, and consequently, its decisional outcomes (i.e., approval, require modifications, or disapproval) for the proposed research.
Relatedly, I was quoted in Science saying that Chat-IRB “could help you handle your more mundane matters, so that you can focus on the substantial stuff”. But what if there might be good reasons for Chat-IRB to assist with, or “do” the substantial?
As such, I believe that when IRBs start to employ GenAI tools in more substantive ways during their ethics review, they begin to enter the zone of ‘AI for moral deliberation’; particularly if and when Chat-IRB is deployed in an agentic fashion to inform and facilitate — and possibly influence in the process — members’ ethical deliberations during their reviews and discussions such as at convened ‘Full Board’ meetings.
Functionally, it is technically feasible to programme a Chat-IRB that provides IRB-users with ethical analyses and justifications for the approval/modification/disapproval of proposed studies. Accordingly, human participation in such AI-assisted ethical deliberations can range — much like in most forms of AI-assisted decision-making — from substantial to little or no meaningful involvement at all.
Unsurprisingly, deploying Chat-IRB in such a manner raises ethical concerns (e.g., automation bias), especially when its outputs are presented or otherwise perceived as highly logical, accurate, and/or authoritative, such that human reviewers begin to anthropomorphise it as an “expert” and accept its recommendations/suggestions at face value.
At the same time, it also raises the more difficult question of whether some degree of deference to a Chat-IRB that demonstrates strong accuracy and reliability in ethical analysis — and perhaps even in moral deliberation — might be warranted, particularly if and when the IRB lacks substantial ethics expertise within its composition. After all, it is quite unlikely that a brief, one-off minimum ethics training course (e.g., CITI program) would equip one with substantial expertise in ethical analysis.
Your proposal assumes past IRB decisions are valuable training data. Has the study team ever considered integrating their tool with the IRB Precedent Tool? If institutions all adopted the IRB Precedent Tool, how would the study team theoretically account for any inconsistencies, biases, or over-conversative IRB decisions?
As described in our paper, Chat-IRB’s functions such as “preliminary review” and “consistency checking” were framed with ‘IRB precedent’ — or perhaps more appropriately termed, ‘decision summary banks’ — in mind.
For transparency, I am a member of the AEREO consortium and am involved in a follow-up project, led by Holly F. Lynch, with its authors/AEREO members examining institutional memory banks for precedent.
Importantly, the introduction of AI should not obfuscate how can and should IRBs operationalise precedent effectively because, the question remains unchanged even in the absence, or non-use, of AI tools; that is to say, how are IRBs to treat banked decisions that are: inconsistent with current proposed decisions; biased; or seemingly over-conservative?
While these are important matters that warrant further analysis and deliberation — and which any precedent project would have to eventually address — here is my brief take:
On inconsistencies: This is why the term ‘decision summary’ is preferable to ‘precedent’ because, the latter might imply or be misunderstood that the IRB’s (future) decisions are strictly bound to past decisions — e.g., for the sake of consistency. This ought to not be the case. Like judges, IRBs should similarly have the ability to deviate from prior decisions if there exist good and justifiable reasons for doing so, and importantly, to disregard prior decisions that are, in retrospect, considered mistaken or made in error.
On bias: Beyond ensuring their decision summaries (DS) are of sufficient quality, IRBs should consider conducting periodic audits of their banked DS to identify decisions that may reflect bias or had been made under biased assumptions.
This is also an area where AI has an advantage over humans. Human bias during reviews is sometimes difficult to detect, and it is even harder to correct. Being told you are biased does not mean you would be able to easily or immediately correct it, nor can there be any assurance that you would cease to be biased in your views and judgements. By contrast, an AI’s outputs can at least be monitored, audited, and iteratively debiased through evaluation and retraining of the model.
Accordingly, when an institution chooses to implement tools such as Chat-IRB, IRB-users and developers should pay close attention to, amongst other things, bias when curating their DS for Chat-IRB’s knowledge base. And if historically biased DS are to be retained in the database to serve as institutional memory, they should at least be flagged accordingly — both to the AI system and to its IRB-users.
On over-conservativeness: To determine whether an IRB decision is or was “overly conservative” in any objective sense would likely require a comparison against some form of baseline. While challenging as it may be, I think establishing a national-level database of sufficiently anonymised DS, contributed by various IRBs, can potentially assist HRPPs in developing a consensus on what constitutes a reasonable standard of human protections (for a given type of study).
That said, it is also important to remember that, as the saying goes, “it depends”. What may seem over-conservative on its face might in fact be appropriate once you consider the contextual factors of the proposed research — e.g., institutional history, cultural considerations, or the prevailing political climate.
Who should be legally and morally accountable when an AI-assisted IRB decision contributes to participant harm?
In terms of legal liability, this would be an important issue for the legal scholars, practitioners, policy makers, and jurisdictions to resolve. However, it is admittedly not an easy question; we can already observe this in analogous applications of AI such as autonomous/self-driving vehicles, where frameworks for determining liability still appear to be under development.
One potential way forward may be to employ “Collective Reflective Equilibrium in Practice”; a method proposed by Savulescu et al. that utilises public preferences/attitudes to help inform and shape policy on the use of controversial novel technologies. (I mention this merely as a possible approach, and not necessarily as an endorsement of it).
At the same time, this also raises a prior question: how many instances have there actually been where an IRB’s decision was determined to have contributed to participant harm? Put differently, how often has harm to a research subject or participant been directly attributed to an IRB’s action or inaction?
Supposing we can establish with certain confidence that the IRB was in fact responsible for the harm caused, then this table by Price et al., which maps “potential physician liability for clinical use of AI” (Table 9.1 from their chapter “Liability for use of artificial intelligence in medicine”), may provide a useful starting point for thinking through the conditions under which an IRB could or would be liable for the harm caused from AI use in its decision-making (see #2 and #7 of the table).
Quite clearly though, this table does not map neatly onto the IRB’s context. The HRPP field will therefore have to develop its own framework for determining IRB lability in relation to the use of AI-assistance for ethics review and decision-making.
Again, these are merely suggestions. As the technology rapidly evolves — and as institutional adoption of AI increases — so too will our legal frameworks and moral analyses change and adapt accordingly.
From my understanding, reasoning models improve explainability because they expose chain-of-thought reasoning. However, there is also AI research that suggests that this approach may be fragile and unreliable. Why should IRBs trust it?
Just as a note on the notion of ‘trust’. There is some philosophical literature on why it would be inappropriate to trust AI in the same way we would trust a human person, and that we should be ‘relying’ on AI instead. Briefly, trust has a relational property (between two agents). So, until AI can be held accountable or responsible for their actions/outputs in ways humans can, it would be more appropriate to ask, “Why should we rely on AI?” rather than, “Why should we trust it?”.
Now, until we possess validated GenAI tools (e.g., LLMs, Reasoners, Agentic AI) that are purpose-built for ethics review, and that achieve accepted, reproducible levels of accuracy and reliability in their outputs and/or Chain-of-Thought (CoT) — similar to the kinds of validation standards expected of for clinical AI — I do not think IRBs should treat a reasoner’s CoT as a panacea for explainability.
That said, this does not mean CoTs are ineffectual or “worthless”. Compared to older models, like the now-retired ChatGPT-4o, which provide end-users with only its final output, a reasoner’s CoT provides at least a certain degree of interpretability regarding how their responses came to be. This permits the human user to assess and determine if the AI might have “hallucinated” its output or made logic errors.
Unsurprisingly, this in turn raises the question of whether the human user possesses the level of expertise required to adequately evaluate the CoT (and the resulting output). Consequently, this relates back to my earlier point regarding appropriate deference to an accurate and reliable AI. What happens if the human user lacks the necessary expertise to critically evaluate the AI’s reasoning and conclusions?
If AI tools in ethics review are indeed inevitable, then it is imperative that institutions begin critically examining the kinds of domain expertise required within their IRBs and their composition, to competently evaluate the outputs of GenAI used in support of, or to augment their ethical judgements, before it is too late.
IRBs are already criticized for potential delays in starting research. Could AI systems unintentionally intensify oversight by generating more flags, more caution, and more procedural scrutiny?
I find statements such as “the IRB delayed my study” to be a somewhat sweeping claim that does not do justice to the many challenges IRBs face in trying to fulfil their mandate.
To be clear, I am not suggesting IRBs are guiltless. But the causes of delays are often nuanced, multifactorial, and sometimes interrelated with knock-on effects, such that it can be unfair to attribute delays to any single point alone. Contributors to extended turnaround times can range from things like:
Sub-optimal staff/member-to-application ratio (i.e., inadequate manpower);
Protracted response times from both researchers and IRBs, especially when the latter are short-handed;
Poorly written submissions, missing documents, and/or discrepancies within the IRB application;
Insufficient knowledge or understanding of ethical, institutional, and/or regulatory requirements;
Poor communication strategies in explaining what is required of researchers — and importantly, the reasons why — owing in part to complex regulatory language; and
Perhaps more controversially, disagreements between the IRB and researchers on what safeguards or restrictions are ethically necessary or justified.
The last point is particularly contentious. IRBs may believe that certain safeguards, alternative procedures, and/or restrictions are necessary to adequately protect participants, or that certain interventions are ethically impermissible because they are too risky. While researchers on the other hand may believe that the IRB’s concerns are unfounded, overly conservative, or insufficiently justified, and that such requirements or changes mandated by the IRB ultimately prevent them from collecting data crucial to answering their research questions.
I believe a well-designed and reliable Chat-IRB would potentially address many of the points raised above. Whether it could help ameliorate or resolve the impasse described in the last point however remains to be seen, though I am cautiously optimistic. For example, it may help IRBs better articulate and justify the rationale underlying their decisions to researchers. Conversely, it could provide IRBs with supporting literature, empirical information, and/or ethical analyses that suggest why their initial assumptions, interpretations, or judgements might be mistaken.
Ultimately, whether Chat-IRB might cause unnecessary delays in securing ethics approval will depend significantly on how, and at which stage of the ethics review process, does the IRB choose to deploy it.
If an IRB determines that delays are arising from AI utilisation, it should then examine the nature of those delays. If the delay stems from the AI identifying ethically (or legally) relevant concerns that the IRB had previously overlooked or insufficiently examined in similar (past) protocols, then such delays would in fact be ethically warranted. In this instance, the AI is helping the IRB better fulfil its societal moral mandate: protecting human research participants, and serving as gatekeepers for the ethical conduct of research.




