Hierarchical document classification, phase 2
Carried the fall research forward — further testing and refinement of the most promising classification approach for Triplan's large document labelling workflow.
During the spring semester, Team Triplan carried out further testing and refinement of the most promising methods from the fall research phase, working towards a solution capable of handling the scale and complexity of Triplan’s classification challenge.
Refining the classification workflow
Triplan’s clients label documents for storage through a large class tree with thousands of leaf nodes — many labels with similar names, making manual sorting confusing and time-consuming. ReLU’s task was to use AI to reduce the number of candidate classes and provide smarter suggestions to end-users.
The team continued extracting text from PDFs without violating personal privacy policies, iterating on the approaches tested in the fall — large language models, classical machine learning, and fine-tuning different types of encoders. The result is a system that meaningfully reduces the burden of manual labeling and brings real value to Triplan’s end-users.