Multimodal
Pocket-Dentist: On-Device Dental Image Understanding via Efficient Multimodal Large Language Models
Pocket-Dentist introduces an efficient benchmark for dental multimodal question answering, utilizing three datasets from BRAR and MetaDent to evaluate 14 vision-language models (VLMs). Notably, a compact 2B-parameter model, Pocket-Dentist-2B, demonstrates competitive performance with larger models while achieving a 4.9x reduction in latency and 2.3x lower memory usage when deployed on an iPhone 17 Pro. This development is significant for practitioners as it enables practical, privacy-preserving dental screening on consumer devices, enhancing accessibility and efficiency in clinical settings.
vision-languagedentalllm