A Hybrid Approach to Object Recognition: From Segmentation to LMM-Based Classification
The goal of the project is to develop an object recognition system based on segmentation components (Segment Anything, SAM) and classification using multimodal models (LMM). This type of system is capable of achieving better results in object recognition for challenging, imbalanced datasets. To date, a proof-of-concept based on a waste sorting case study (the TACO and WaRP datasets) has been completed. The experimental results (classification accuracy on a subset of images of 80% for GPT-4o and 30% for Llama3.2-Vision), presented at PP-RAI 2025, were deemed promising. The activities planned by the applicants include: developing a module to filter objects returned by SAM (to improve system performance), expanding the experiments to include full test sets of reference datasets, and improving the performance of Llama3.2-Vision (prompt engineering, possibly fine-tuning). In a subsequent phase, experiments with other LMM models, expansion of the dataset collection, and integration of an explainability component are planned.
Machine-translated
Przetłumaczone maszynowo