e2tree is built by researchers at the University of Naples Federico II and K-Synth. The goal: make ensemble models auditable without sacrificing accuracy.
Authors of the e2tree method and maintainers of the R package.
Co-authors on the papers that introduced and extended the method.
Three papers describe the method. The first covers classification, the second extends it to regression, and the third builds the statistical framework for measuring how faithfully the tree reconstructs its ensemble.
Aria, M., Gnasso, A., Iorio, C., & Pandolfo, G.
Introduces the algorithm for classification. Defines the co-occurrence dissimilarity framework and the recursive tree-growing procedure, benchmarked on real datasets. The single e2tree matches Random Forest accuracy and is readable.
Aria, M., Gnasso, A., Iorio, C., & Fokkema, M.
Extends the method to regression. Co-occurrence weighting is adapted to account for response similarity between pairs, with a new impurity measure for continuous targets. Benchmarked against CART and GUIDE.
Aria, M., Gnasso, A., & Iorio, C.
Validating a surrogate requires measuring agreement with the ensemble's internal representation, not mere association. This paper introduces the normalized Loss of Interpretability (nLoI), a robustified Neyman-type statistic sitting in the Cressie–Read power divergence family at λ = −2, whose closed-form decomposition into within- and between-node components identifies precisely where and why reconstruction fails. Four complementary measures — Hellinger, wRMSE, RV coefficient and SSIM — and a unified permutation testing procedure complete the toolkit.
Submitted to Computational Statistics & Data Analysis. Preprint available on request — implemented in the package as loi(), loi_perm() and localLoI().
Cite the relevant paper and the R package: