ADD. Revista Latinoamericana de Psicología · 2026
Evaluating early development with AI: A multimodal analysis of caregiver–toddler interactions in Argentina
The reliable assessment of complex caregiver–toddler behaviours continues to be a core challenge in developmental psychology. This study presents a methodological proof-of-concept comparing human evaluators and specific proprietary AI models using a unified, rubric-based multimodal protocol. We analyse how successive iterations of Gemini (1.5, 2.0, 2.5), individual experts, and collaborative student groups interpret and score naturalistic dyadic interactions in Argentina. We note that this asymmetric design—student consensus versus individual experts—confounds rater identity with the evaluation process, and the small sample size restricts statistical power. Method: Five rater groups evaluated videos using a unified 35-item rubric. We combined classical indices (exact percent agreement, weighted kappa, intraclass correlation coefficient [ICC]) with a hierarchical Bayesian model to estimate rater internal variability and latent video scores. Results: Under the tested priors, the most recent AI (Gemini 2.5) exhibits internal scoring variability comparable to individual experts. However, high precision does not imply high agreement: AI’s scoring pattern departs from experts, yielding low ICC despite moderate exact agreement, indicating that AI is not interchangeable with human judgment. Notably, student groups operating by consensus exhibit the highest precision, illustrating a robust form of ’social wisdom’. Conclusions: The findings highlight the rapid evolution of AI capabilities but underscore the distinct nature of its evaluative patterns compared to humans. We discuss implications for prompt engineering, cultural–linguistic specificity, and the fast pace of model evolution, cautioning that deployment in LMIC settings requires rigorous validation guardrails rather than assumed expert-equivalence.