Document resource
Introduction The field of Risk Assessment is challenged by the large quantity of literature it needs to process: a huge labor-intensive task. Automated risk-assessment aims to support researchers by providing tools to extract large quantities of information. One of these tools is INDRA (Integrated Network and Dynamical Reasoning Assembler), which makes use of several reading systems (i.e. we selected five according to developer’s scale of usage and efficiency: REACH, EIDOS, SPARSER, ISI, TRIPS), to extract mechanistic information from texts. These are captured in statements, a standardized output format. Currently, accuracy and coverage of the independent reading systems are not well-established. Therefore, in the current study we assess those properties of these reading systems, and evaluated how data-fusion can be applied to further increase their reliabilities.Material and Methods Each reader was run on a manually curated cancer-related ground-truth subset of 12 full-text articles selected for mechanistic richness and variety of content. Next, the set of extracted statements was reviewed for errors of any type (e.g. from entity type- such as proteins, to directions). Moreover, coverage was assessed. Lastly, Random Forest Modelling and Data Fusion were performed to predict errors.Results The current pipeline consolidated 1860 mechanistic statements, which have been matched with 1376 manually extracted statements from the same papers. Generally, we observed that the highest number of statements comes from EIDOS (c.a 1000) and REACH (c.a 600). Reach showed lowest false-positive rate. More specific results (i.e. per reading system, per entity type) will be shown.Conclusion It is crucial to establish the reliability of the INDRA tool to interpret results and enable improvements to the algorithm. The current implementation will be further examined in Lexces2, and the European projects of Endomix and EDC-MASLD.