ORIGINAL
ISSN 0120-0534
Volumen 58
* Corresponding author.
E-mail: jose.amorocho@unisabana.edu.co https://doi.org/10.14349/rlp.2026.v58.2 0120-0534/© 2026 Fundación Universitaria Konrad Lorenz. This is an open access article under the CC BY-NC-ND license (https://creativecommons.org/licenses/by-nc-nd/4.0/).
Revista Latinoamericana de Psicología (2026) 58, e580277 https://doi.org/10.14349/rlp.2026.v58.2 https://revistalatinoamericanadepsicologia.konradlorenz.edu.co/ Evaluating early development with AI:
A multimodal analysis of caregiver–toddler interactions in Argentina José Amorocho a,* , Juan José Giraldo-Huertas a , Lucas G. Gago-Galvagno b,c
,
Angel M. Elgier b,c , Natalia A. Mancini b,c
a Universidad de la Sabana, Colombia b Universidad Abierta Interamericana, Argentina c Universidad de Buenos Aires, CONICET, Argentina Received 23 October 2025; accepted 13 April 2026 Abstract | Objective: The reliable assessment of complex caregiver–toddler behaviours continues to be a core challenge in developmental psychology. This study presents a methodological proof-of-concept comparing human evaluators and specific proprietary AI models using a unified, rubric-based multimodal protocol. We analyse how successive iterations of Gemini (1.5, 2.0, 2.5), individual experts, and collaborative student groups interpret and score naturalistic dyadic in teractions in Argentina. We note that this asymmetric design—student consensus versus individual experts—confounds rater identity with the evaluation process, and the small sample size restricts statistical power. Method: Five rater groups evaluated videos using a unified 35-item rubric. We combined classical indices (exact percent agreement, weighted kappa, intraclass correlation coefficient [ICC]) with a hierarchical Bayesian model to estimate rater internal variability and latent video scores. Results: Under the tested priors, the most recent AI (Gemini 2.5) exhibits internal scoring variability com parable to individual experts. However, high precision does not imply high agreement: AI’s scoring pattern departs from experts, yielding low ICC despite moderate exact agreement, indicating that AI is not interchangeable with human judg ment. Notably, student groups operating by consensus exhibit the highest precision, illustrating a robust form of ’social wisdom’. Conclusions: The findings highlight the rapid evolution of AI capabilities but underscore the distinct nature of its evaluative patterns compared to humans. We discuss implications for prompt engineering, cultural–linguistic specificity, and the fast pace of model evolution, cautioning that deployment in LMIC settings requires rigorous validation guardrails rather than assumed expert-equivalence.
Keywords: Human-computer interaction, psychological assessment, artificial intelligence, rater reliability, developmen tal psychology, Bayesian modelling, scaffolding © 2026 Fundación Universitaria Konrad Lorenz. This is an open access article under the CC BY-NC-ND license (https://creativecommons.org/licenses/ by-nc-nd/4.0/).
Evaluación del desarrollo temprano con IA: análisis multimodal de interacciones cuidador–niño en Argentina Resumen | Introducción/objetivo: La evaluación fiable de las complejas conductas entre cuidadores y niños pequeños es un desafío central en la psicología del desarrollo. Este estudio presenta una prueba de concepto metodológica comparando evaluadores humanos y modelos específicos de IA patentada utilizando un protocolo multimodal unificado. Analizamos cómo versiones sucesivas de Gemini (1.5, 2.0, 2.5), expertos individuales y grupos colaborativos de estudiantes interpretan
2/11 J. Amorocho et al.
Artificial intelligence in behavioural assessment Artificial intelligence (AI) is a field of computer science that focuses on creating systems or machines that per form automatic tasks (e.g., recognising patterns, inte racting with environments, learning from experience) that normally require human abilities (Pelau et al., 2021). Recent advances, spurred by the influential Transfor mer architecture (Vaswani et al., 2017), have culminated in sophisticated multimodal models such as GPT and Gemini. The core innovation of these models is the at tention mechanism, a computational method analogous to cognitive focus, which allows the model to dynami cally weigh the importance of different segments of input data when generating a response.
In a multimodal context, this concept is extended to cross-modal attention, enabling the system to learn di rect correlations between different data streams. For example, it can map specific visual features in a video (e.g., a child’s pointing gesture) to corresponding lin guistic cues in the audio (e.g., the caregiver naming the object). This capacity for deeply integrating visual and linguistic input is foundational, evolving from the prin ciples pioneered in seminal language-only models such as GPT (Radford et al., 2018) and BERT (Devlin et al., 2018). These systems, part of the broader field of Generati ve Artificial Intelligence (GenAI), have approached hu man performance on specific generation and unders tanding benchmarks (Brown et al., 2020; Goodfellow et al., 2014). Consequently, early work already demonstra tes their potential for the automated scoring of com plex behavioural interactions (Nagrani et al., 2021; Sha fiee Rad, 2025; Su & Yang, 2022). Recent studies focusing specifically on Large Language Models (LLMs) propose calibrated rubric frameworks for evaluation (Hashemi et al., 2025; Pathak et al., 2025) and highlight differences between LLMs and human graders in scoring processes (Chu et al., 2024; Wu et al., 2025).
y califican interacciones didácticas naturalísticas en Argentina. Notamos que este diseño asimétrico —consenso de estu diantes frente a expertos individuales— confunde la identidad del evaluador con el proceso de evaluación, y el pequeño tamaño de la muestra restringe el poder estadístico. Método: Cinco grupos de evaluadores calificaron videos utilizando una rúbrica unificada de 35 ítems. Se combinaron índices clásicos (porcentaje de acuerdo exacto, kappa ponderado, co eficiente de correlación intraclase [CCI]) con un modelo jerárquico bayesiano para estimar la variabilidad interna de los evaluadores y las puntuaciones latentes de los videos. Resultados: Bajo los priors evaluados, la IA más reciente (Gemini 2.5) alcanzó una variabilidad interna de puntuación comparable a la de los expertos; sin embargo, una alta precisión no impli ca un alto acuerdo: el patrón de calificación de la IA difirió del de los expertos, lo que resultó en un CCI bajo a pesar de un acuerdo exacto moderado, indicando que la IA no es intercambiable con el juicio humano. De forma destacada, los grupos de estudiantes que evaluaron por consenso lograron la mayor precisión, demostrando un efecto de ’sabiduría colectiva’. Conclusiones: Los hallazgos destacan la rápida evolución de las capacidades de la IA, pero subrayan la naturaleza distinta de sus patrones evaluativos en comparación con los humanos. Este artículo ofrece un marco para el despliegue riguroso y ético de herramientas de evaluación con IA en países de ingresos bajos y medios, advirtiendo que su despliegue requiere rigurosas barreras de validación en lugar de una equivalencia experta asumida. Palabras clave: Interacción humano-computador, evaluación psicológica, inteligencia artificial, fiabilidad entre evaluado res, psicología del desarrollo, modelado bayesiano, andamiaje © 2026 Fundación Universitaria Konrad Lorenz. Este es un artículo Open Access bajo la licencia CC BY-NC-ND (https://creativecommons.org/licenses/ by-nc-nd/4.0/).
Scaffolding and intersubjectivity in early development Despite these technological advances, applying them to nuanced psychological phenomena presents signi ficant challenges. Within the sociocultural framework of Vygotsky (Vygotsky, 1978), learning occurs through guided interaction within the Zone of Proximal Deve lopment, where more knowledgeable partners scaffold a learner’s emerging capabilities (Bruner, 1978). Scaffol ding refers to the process by which caregivers provide tailored, temporary support such as contingent ques tioning, modelling, and providing feedback to guide a child through a task. The quality of these interactive behaviours is a strong predictor of later cognitive and literacy outcomes (Dickinson & Porche, 2011; Mol & Bus,
2011).
A powerful application of these principles is seen in Dialogical Book Sharing (DBS), which extends sca ffolding to shared reading contexts. DBS emphasises reciprocal, inter-subjective behaviours such as gaze-fo llowing, deictic gestures, and emotional attunement that are known to enrich language growth in young children (Murray et al., 2022; Vally, 2012). Empirical stu dies across diverse settings have confirmed the efficacy of DBS, showing that training caregivers in these dia logic techniques yields significant gains in vocabulary, narrative comprehension, and attention, particularly in children from low-resource communities (Balog et al., 2024; Hurtado-Mazeyra et al., 2024; Koopowitz et al., 2024; Murray et al., 2023; Vally et al., 2015). However, reliably measuring these intersubjective scaffolding behaviours remains a central challenge in developmental psychology. Human coding is not only time-consuming and costly but also subject to signifi cant rater variability and bias (Carballo-Fazanes et al., 2021; Pinheiro-Carozzo & Murta, 2023; Pontes & Brino, 2022; Wind, 2019). In contexts such as Colombia’s Sabana Centro region, where parent–child reading interactions often occur under informal, naturalistic conditions (Gi
3/11 Evaluating early development with AI: A multimodal analysis of caregiver–toddler interactions in Argentina raldo-Huertas et al., 2023), the need for robust, automa ted tools capable of handling real-world video and au dio data becomes even more pressing.
The assessment challenge in lowand middle-income countries The challenges of reliable and accessible psychologi cal assessment are exacerbated in the context of Latin America (LATAM) and other Low- and Middle-Income Countries (LMICs).
Educational inequities in these regions remain stark, creating a digital divide that undermines early childhood development (Orozco Restrepo et al., 2022). These systemic issues are rooted in profound socioe conomic realities. The levels of poverty in the region are about 27.3 percent, which translates to 172 million indi viduals living in low-SES contexts, and extreme pover ty affects about 10.6 percent of the population (CEPAL, 2024). In this scenario, scalable and low-cost automated evaluation methods particularly those leveraging ubi quitous mobile technologies and AI offer a promising means to democratise access to formative assessment and support equity in developmental interventions (Motiwalla, 2007; World Bank, 2021).
The present study Accordingly, the present study addresses the intersec tion of these needs and opportunities. While the im portance of scaffolding is well-established and the po tential of GenAI is clear, direct comparisons between successive AI versions and human raters on the same interactional data are scarce. While Temple (2024)’s pi lot thesis compared LLM outputs to student coders, it lacked multimodal video inputs. Furthermore, Goert zen and Klaus (2023) highlight the critical importance of using rigorous interrater reliability metrics (e.g., ICC, percent agreement) to properly validate and contextua lise any automated assessment. To our knowledge, no study has yet directly compared Gemini 1.5 with Gemi ni 2.0, Gemini 2.5, undergraduate psychology students, and experienced experts in a unified evaluation of the se critical parent–child scaffolding behaviours. Concu rrently, LATAM studies emphasise the centrality of di rect observation and shared reading/play as proximal targets for early development (Balog et al., 2024; Hur tado-Mazeyra et al., 2024; Pinheiro-Carozzo & Murta, 2023; Pontes & Brino, 2022).
We aim to address this gap by comparing ratings from these distinct evaluator groups, employing both classical interrater indices (exact agreement, weigh ted ICC; (Goertzen & Klaus, 2023; Hallgren, 2012)) and a hierarchical Bayesian model to estimate each rater’s internal variability scoring pattern and the latent “true” score for each video. Contextualising our work, previous studies comparing AI to human assessment have yielded varied inter-rater agreement. While some machine learning models achieve performance com parable to human raters (Carcone et al., 2019), reliabi lity can vary widely. Similarly, inter-rater agreement among human coders for complex behaviours ranges from low for novices (Oremus et al., 2012) to excellent for trained experts using established tools (Montirosso et al., 2023; Vilaseca et al., 2019), underscoring the need for direct comparisons against multiple human bench marks. Crucially, our study maintains the procedural asymmetry between student consensus groups and in dividual experts; this limits direct rater-to-rater equi valent comparisons but reflects feasible deployment workflows. Our primary research question is: How ac curately do specific proprietary multimodal AI models (Gemini) evaluate scaffolding interactions compared to human raters? A secondary question explores whe ther performance improves across AI versions and how AI-generated internal consistency compares to that of novice and expert humans.
Finally, in light of the rapid evolution of multimodal AI systems (e.g., Anil et al., 2024; Wu et al., 2025), we fra me the present work as a methodological proof-of-concept for comparing human and AI evaluators under a com mon rubric, bounded by a specific, closed-weight model family (Gemini) due to the lack of open-model baseline comparisons. This perspective highlights design choi ces that matter for validity such as prompt engineering and the selection of reliability metrics while recogni sing that conclusions must be revisited as models im prove at a fast pace (Radanliev, 2024; Wang et al., 2024). By grounding this investigation in a middle-income context with real-world video data, this study aims to inform scalable assessment solutions that promote equity in early childhood interventions. Method This exploratory instrumental study compared five rater conditions-Gemini 1.5, Gemini 2.0, Gemini 2.5, undergraduate psychology students, and developmen tal-psychology experts-scoring the same set of 14 care giver–child videos with an identical 35-item rubric (1–9 scale; frequency, quality, impact). We contrasted AI and human ratings using classical interrater indices (Per cent Agreement, Weighted Kappa, ICC) and a hierarchi cal Bayesian model (PyMC) to estimate each rater’s pre cision and the latent video scores.
Participants and raters Video corpus. Fourteen smartphone recordings (mean length ≈ 8 minutes) captured unstructured caregiver– child play in Argentina under naturalistic conditions (no strict controls on resolution or audio). Written in formed consent was obtained; procedures respected privacy and human subjects protections. The corpus size enabled intensive, multi-rater comparison on in formation-rich cases (Guest et al., 2006; Hertzog, 2008). Student raters. Thirty undergraduates (Mage = 20.1, SD = 1.4; 77% female) rated in pairs/triads after brief tra ining on the rubric foundations (E.S.A., DBS). Expert raters. Three developmental-psychology ex perts (2 male, 1 female; Mage = 30, SD = 1), all Ph.D. and > 5 years of experience, rated independently. AI raters. Gemini 1.5, Gemini 2.0 (Flash), and Gemini
2.5 (Flash-preview) received identical prompts and pre
processing; only the underlying model version varied.
4/11 J. Amorocho et al.
Inferences were obtained via the vendor API; downs tream analyses used PyMC.
Materials Rubric. The 35-item instrument integrates the Esca la de Sensibilidad del Adulto (Adult Sensitivity Sca le -E.S.A.-; Santelices et al., 2012) with Dialogical Book Sharing (DBS) criteria (Murray et al., 2022), yielding a unified space for affective attunement, responsiveness, scaffolding, and cognitively supportive talk. Items were organised into three higher-order dimensions (empa thetic response, playful interaction, emotional expres sion) with harmonised anchors; each item was scored 1–9 based on observed frequency, quality, and impact (Norrie et al., 2024). These constructs overlap with ob servational schemes recently applied in LATAM clinical and family contexts, supporting content validity for our coding space (Pinheiro-Carozzo & Murta, 2023; Pon tes & Brino, 2022).
Automated pipeline. Videos were normalised to MP4, segmented into 5-second clips, and processed via automatic speech recognition (ASR) and computer-vi sion (CV) modules using open-source libraries (Nagrani et al., 2021; Vally et al., 2015). The ASR component trans cribed caregiver and child speech; based on published benchmarks for comparable Spanish-language, natu ralistic audio conditions, word error rates are typically in the range of 20–40% (Baevski et al., 2020; Radford et al., 2023), a range consistent with the informal recor ding quality of our corpus. The CV component extracted frame-level features including facial-expression valen ce estimates and gaze direction; published accuracy figures for such features under real-world, low-control conditions are generally moderate (e.g., AUC ≈ 0.70–0.80 for facial action unit detection; Kollias et al., 2023). The generated segment descriptions were then scored by the Gemini models against the rubric. Due to the proofof-concept nature of this study, exhaustive formal qua lity-control metrics for these preprocessing steps (e.g., Word Error Rate, frame-drop rates) and single-modali ty ablation checks were not systematically documented for our specific corpus, which remains a limitation for assessing cascading pipeline errors. Future iterations should benchmark ASR and CV performance directly on the target corpus prior to model scoring. Coverage and targets. Analyses used item-level tar gets defined as video×item. Due to quality control (e.g., missing audio segments, ASR/vision filters) and the re quirement that pairs of raters share observations on the same targets, the number of pairwise targets (n) va ries across rater pairs (reported in Table 2). Procedure Human rating. Students first offered individual im pressions and then briefly discussed discrepancies to reach a consensus score (functioning like a small panel that averages idiosyncrasies). Experts viewed each vi deo once and rated independently (no deliberation). This methodological asymmetry confounds rater identity with the process-of-rating; thus, comparisons between students and experts reflect differences in both exper tise and consensus-building workflows.
AI rating. The same preprocessed clips and rubric prompt were submitted to each Gemini version with controlled generation settings favouring consistency: temperature = 0.3, top_p = 0.8, top_k = 32, max output tokens = 500 (Hu et al., 2023).
Duration and ethics. Experts completed ratings over 2 weeks; student groups in a single 2-hour ses sion. On a single machine, AI scoring took 8 minutes/ video for the scoring pass (excluding normalisation and QC). The protocol was approved by the Ethics Commit tee of the Faculty of Behavioural Science, Universidad de La Sabana (Act No. 200, February 21, 2024), and ad hered to the Declaration of Helsinki and American Psy chological Association ethical principles. The protocol was deemed minimal risk. Written informed consent was obtained from legal guardians and assent from children when appropriate for voluntary participation. Participant confidentiality was maintained through de-identification, pseudonymisation, and secure data storage. For automated analysis, only de-identified tex tual descriptions were submitted to external APIs with data retention disabled. AI-generated outputs served exclusively for methodological comparison under hu man oversight and did not inform any clinical deci sions regarding participants. Data sharing is restricted to de-identified aggregate data under a formal data-use agreement.
Data analysis Classical interrater indices. Data were analysed in a long format respecting the item-wise ordinal struc ture without mean-imputation across subscales. We computed Percent Agreement (PA) and ICC (3,1) (twoway mixed, single-measure, consistency; raters trea ted as fixed). PA was defined as the proportion of exact item-level matches and adjacent-category matches (PA ±1) calculated strictly over shared video-item targets to account for discrepancies in coverage n across pairs. To address chance alignment and ordinal structure, we additionally computed linearly weighted Cohen’s Kappa. ICC (3,1) point estimates were obtained via twoway ANOVA components, and 95% confidence intervals were estimated via percentile bootstrap (2,000 resam ples) by resampling targets (video × item) within each rater pair. Because pairwise ICCs involve k = 2 raters per comparison and variable item coverage across pairs, bootstrap Cis provide more robust uncertainty quanti fication than closed-form approximations. Bayesian hierarchical model. To estimate each ra ter’s precision and the latent “true” score for each video, we specified in PyMC:
xrvj N(µv, sr), where xrvj is the score by rater r on video v, item j. Priors µv& N(7, 2), v = 1,...,14, sexpert& Half-Cauchy(0,0.5),
5/11 Evaluating early development with AI: A multimodal analysis of caregiver–toddler interactions in Argentina sstudent& Half-Cauchy(0,1.0), sGemini2.5 Half-Cauchy(0,1.25), sGemini2& Half-Cauchy(0,1.5), sGemini1& Half-Cauchy(0,2.0).
These priors reflect expected relative precision from the literature (experts > students > earlier AI), while allowing uncertainty about LLM evaluation tasks (e.g., Anil et al., 2024; Brown et al., 2020; Lin et al., 2024; Wind, 2019). Posterior inference used MCMC (4 chains, 2,000 draws, 1,000 tune); rater precisions were compared via posteriors of sr −2.
Post-hoc sensitivity. With 14 videos and unequal group sizes, simulation checks indicated power prima rily for very large effects (Cohen’s d ≈ 1.9); reproducible code is provided in the companion Colab notebook. We emphasise that post hoc power is not appropriate for confirming null effects; rather, we report sensitivity to indicate the smallest effects our design could reasona bly detect under standard assumptions. Full simulation code and sensitivity reports are provided in the repro ducible notebook.
Results Descriptive statistics Score distributions differed across raters (Figure 1). Hu man Student and Human Expert showed comparable central shapes, with Experts slightly more dispersed towards higher scores. Gemini 2.0 occupied a range si milar to humans but with a broader spread; Gemini 1.5 was diffuse at lower–mid values, and Gemini 2.5 clus tered at the upper end. Table 1 quantifies these patter ns: in this sample, Gemini 2.5 had the highest mean (M = 8.83) and lowest variability (SD = 1.14) ; Gemini 1.5 the lowest mean (M = 4.63) and highest variability (SD = 3.53). Human Experts averaged higher (M = 7.38, SD = 1.96) than Students (M = 7.02, SD = 1.87). Gemini 2.0’s mean (M = 6.76) approximated Students but with a subs tantially greater spread (SD = 3.34).
Table 1. Descriptive statistics for rubric scores by rater Rater M SD n Gemini 1.5
4.63
3.53
429
Gemini 2.0
6.76
3.34
396
Gemini 2.5
8.83
1.14
396
Human Expert
7.38
1.96
396
Human Student
7.02
1.87
289
Classical interrater reliability Exact Percent Agreement (PAexact) ranged from 18.2% to 60.9%, tolerant agreement (PA±1) from 28.0% to 62.1%, and ICC(3,1) from −0.40 to 0.17 (Table 2). Gemini 2.0 vs. Human Expert showed the highest ICC(3,1) (0.17; 95% CI [0.07, 0.27]) with PAexact = 27.8%, PA±1 = 54.5%, and wei ghted k = 0.13. Gemini 2.0 vs. Gemini 2.5 exhibited the highest PAexact (60.9%) yet a negative ICC(3,1) = −0.08 [−0.17, 0.03] and low weighted k = 0.06, illustrating that frequent exact matches need not imply high correla tion-based reliability when scale variance is used di fferently across raters. Strongly negative ICCs (Gemini
1.5 vs. Gemini 2.5: −0.40 [−0.46, −0.33]; Gemini 2.5 vs. Hu
man Expert: −0.24 [−0.28, −0.19]) reflect the systematic ceiling clustering of Gemini 2.5 against the broad dis persion of other raters.
Bayesian hierarchical model The model converged (Rˆ = 1.00 for all parameters; effec tive sample sizes > 7,000). Latent video scores µv varied Figure 1. Violin plot of score distributions by rater Gemini 2.0 Human Student Gemini 1.5 Human Expert Gemini 2.5
1
2
3
4
5
6
7
8
9
Score Score Distribution by Rater
6/11 J. Amorocho et al.
meaningfully (Figure 2): video 4 was the lowest (M = 6.57,
95% HDI [6.30,
6.84]) and video 7 was the highest (M = 8.43, 95% HDI [8.16, 8.72]); overall, µv means ranged 6.57–8.43. Precision parameters (sr; lower is more precise) distinguished rater groups (Figure 3): Gemini 1.5 s = 4.52 [4.21, 4.82]; Gemini 2.0 s = 3.16 [2.95, 3.38]; Gemini 2.5 s = 1.81 [1.66, 1.95]; Human Expert s = 1.80 [1.66, 1.93]; Human Student s = 1.37 [1.29, 1.46]. Thus, the Student consensus was the most precise, the Experts and Gemini 2.5 were com parable, and earlier Gemini versions were less precise, demonstrating that high precision scoring can emerge independently of structural agreement.
These conclusions held under less-informative and unit-scale priors; diagnostics were satisfactory across specifications (Rˆ ≈ 1.00, high bulk/tail ESS). Discussion This study compared successive generative AI models (Gemini 1.5/2.0/2.5) with human raters (student con sensus groups and individual experts) when scoring caregiver–toddler scaffolding. Using a unified rubric on identical videos, we combined classical interrater indi ces with a hierarchical Bayesian model to distinguish precision (internal consistency of a rater) from agreement Table 2. Classical interrater reliability indices for rater pairs (sorted by ICC), computed on strict pairwise overlapping tar gets without imputation Rater Pair PA_exact (%) PA±1 (%) Wt. k
ICC(3,1)
95% CI n Gemini 2.0 vs. Human Expert
27.8
54.5
0.13
0.17
[0.07, 0.27]
396
Gemini 2.0 vs. Human Student
24.0
43.1
0.10
0.13
[0.01, 0.24]
204
Human Expert vs. Human Student
25.0
52.5
0.10
0.08
[−0.05, 0.20]
204
Gemini 1.5 vs. Gemini 2.0
32.1
36.4
0.12
0.04
[−0.07, 0.15]
330
Gemini 1.5 vs. Human Student
18.2
35.9
0.10
−0.06 [−0.20, 0.07]
170
Gemini 2.0 vs. Gemini 2.5
60.9
62.1
0.06
−0.08 [−0.17, 0.03]
330
Gemini 2.5 vs. Human Student
27.6
53.5
0.03
−0.11 [−0.27, 0.05]
170
Gemini 2.5 vs. Human Expert
31.8
59.4
−0.02 −0.24 [−0.28, −0.19]
330
Gemini 1.5 vs. Human Expert
19.4
35.8
−0.04 −0.34 [−0.43, −0.25]
330
Gemini 1.5 vs. Gemini 2.5
26.3
28.0
0.02
−0.40 [−0.46, −0.33]
396
Note. PA_exact = exact Percent Agreement; PA±1 = tolerant PA allowing adjacent-category matches (±1 point). Wt. k = linearly weighted Cohen’s Kappa for ordinal scale 1–9. ICC(3,1) = two-way mixed, single-measure consistency; 95% CI via percentile bootstrap (2,000 resamples). n = shared video×item targets per pair. All metrics computed in strict long format without imputation. (alignment with others). The findings provide an initial snapshot of multimodal AI performance in real-world behavioural assessment and specify conditions for res ponsible use in resource-constrained settings. Central findings and alignment with prior work Three results stand out. First, there was clear genera tional improvement across the Gemini series: Gemini
1.5 showed the lowest means and the greatest variabili
ty, Gemini 2.0 was intermediate, and Gemini 2.5 yielded the highest means with the tightest spread, indicating increasing stability with model iteration. The Bayesian analysis paralleled these patterns in rater precision (sr): Gemini 1.5 was the least precise, Gemini 2.0 improved, and Gemini 2.5 achieved the highest precision (lowest sr) of any individual rater, surpassing even Human Ex perts, though this tight clustering occurred at the cei ling of the rating, with distinct alignment patterns, consistent with reports of rapidly advancing multimo dal capabilities (Wu et al., 2025). The capacity of Gemini
2.5 to achieve levels of rater precision (in the Bayesian
sense) comparable to individual human experts echoes findings in other domains where AI has demonstrated Figure 2. Trace and posterior distribution for latent video-level scores µv. (a) Trace plot across MCMC iterations. (b) Poste rior densities for each video’s latent score
6.0
6.5
7.0
7.5
8.0
8.5
9.0
mu_video
0
250
500
750
1000
1250
1500
1750
6
7
8
9
mu_video
7/11 Evaluating early development with AI: A multimodal analysis of caregiver–toddler interactions in Argentina Figure 3. Trace and posterior distribution for rater precision parameters sr. (a) Trace plots across MCMC iterations. (b) Pos terior densities (lower values denote higher precision)
0
1000 2000 3000 4000 5000 6000 7000 8000
Sample
1.5
2.0
2.5
3.0
3.5
4.0
4.5
5.0
σ Sigma trace by rater
1.5
2.0
2.5
3.0
3.5
4.0
4.5
5.0
σ
0
2
4
6
8
Density Posterior of sigma by rater Source Gemini 2.0 Human Student Gemini 1.5 Human Expert Gemini 2.5 performance on a par with or exceeding humans on specific metrics (Brown et al., 2020; Nagrani et al., 2021). Second, precision did not imply agreement. Despite moderate PAexact (= 31.8%), Gemini 2.5 showed a strongly negative ICC(3,1) = −0.24 [−0.28, −0.19] with individual Experts, underscoring that an AI can be internally con sistent yet follow a different scoring behaviour, charac terised by ceiling effects, relative to experts. This dis sociation echoes prior work showing that LLMs can be reliable by some statistics while diverging from human graders in variance use and decision profiles (Goertzen & Klaus, 2023; Hashemi et al., 2025; Wu et al., 2025). The classical indices further illustrate this point: Gemini
2.0 vs. Gemini 2.5 achieved the highest PAexact (60.9%)
but a negative ICC(−0.08) and weighted k = 0.06. Such mismatches arise when raters use scale variance diffe rently—here, we observed upper-end compression in Gemini 2.5 (ceiling-leaning profiles) versus mid/low clustering in earlier models—reducing covariance even when exact matches occur at common anchors (Goert zen & Klaus, 2023; Wu et al., 2025). Under our prompting and preprocessing constraints, Gemini 2.5 tended to as sign higher anchors, which may compress upper-scale variance relative to experts.
Third, student consensus groups were the most pre cise raters in the Bayesian model, exceeding both indi vidual Experts and all AI versions. This likely reflects a “social wisdom” effect: brief, structured deliberation attenuates idiosyncratic error and stabilises rubric use (Pike, 2005). Their mean levels were close to Gemini 2.0 and slightly lower than Experts, with generally mo dest ICCs vis-à-vis other raters. Methodologically, this highlights that process matters: not only who rates, but how. This asymmetry confounds the effect of evalua tor expertise (novice vs. expert) with the effect of the evaluation procedure (consensus vs. individual). Con sequently, direct comparisons of their scores as pure ly measures of individual rater acumen versus expert acumen should be made with caution; rather, they re flect the output of these distinct processes. This also aligns with evidence that trained students can be re liable under structured protocols, while still differing from experts on patterning (cf. Temple, 2024). Context may further shape alignment. Our Spani sh-language, Argentinian corpus could expose cultu ral-pragmatic cues (e.g., prosodic warmth, deictics, re gionally marked child-directed speech) underweighted by models trained largely on English-centric data, mo tivating culturally localised calibration or fine-tuning and cross-site validation. More broadly, prior work finds that machine learning systems can approach human performance yet diverge in agreement and error profi les depending on task and instrument (Carcone et al., 2019; Montirosso et al., 2023; Oremus et al., 2012; Vilaseca et al., 2019). Together, these results suggest treating AI precision and human alignment as distinct targets in developmental assessment pipelines.
Implications for practice, policy, and ethics (with a focus on LMICs) Taken together, the profile observed for Gemini 2.5 -pre cision comparable to experts on this dataset, with distinct alignment patterns - positions GenAI as a collaborator, not a replacement, for scalable assessment in Low- and Middle-Income Countries (LMICs). In settings with few specialists, models can process smartphone video to ge nerate rubric-based summaries, triage cases for human review, and provide consistent baselines for program me monitoring freeing expert time for complex judg
8/11 J. Amorocho et al.
ments, capacity-building, and targeted interventions (Idrovo, 2024; Marko et al., 2025; Van Noorden & Perkel, 2023). In LATAM, scalable measurement can specifically track gains from shared reading and daily parent–child play programmes (Balog et al., 2024; Hurtado-Mazeyra et al., 2024). Because capabilities evolve quickly, deploy ments should treat each model update as an intervention requiring version control (frozen snapshots, change logs, regression tests) and periodic re-validation against fixed benchmarks (Radanliev, 2024; Wang et al., 2024). A concise ethical posture is essential for LMIC use: (i) human-in-the-loop oversight-AI supports, clinicians decide; (ii) robust consent explaining probabilistic ou tputs and intended uses; (iii) equity via local calibration and monitoring for cultural/linguistic bias; (iv) trans parency/contestability through interpretable sum maries; and (v) data stewardship proportional to risk (de-identification, secure handling; Hashemi et al., 2025; Johnson & Stead, 2022; Ong et al., 2025; Saab et al., 2024; Sahdra et al., 2024; Wu et al., 2025). These safeguards mitigate risks of cultural misalignment, performance drift, and overreliance on single-model judgments (Xa mes, 2025) while preserving local validity and trust. Limitations and future directions Interpretation should be cautious. The corpus compri sed 14 naturalistic caregiver–child play videos from Argentina; post-hoc sensitivity indicated power pri marily for very large effects (Cohen’s d ≈ 1.89), so sma ller between-rater differences may be undetected. Only three Experts participated, limiting inferences about expert variability. Human procedures differed by de sign: Students produced consensus scores after training, whereas Experts rated individually; outputs therefore reflect distinct processes as well as individual acumen. Although prompts and sampling were standardised, modest wording shifts can steer model scoring. Finally, we relied on a single integrated rubric, limiting instru ment generalisability.
Three priorities follow.
First, cross-task and cross-language validation: test robustness across in teraction types (play, teaching, problem solving) and languages/dialects to probe cultural transportability and cue weighting, and cross-validate with alternati ve instruments, including family shared reading and free-play tasks documented in recent LATAM studies (Balog et al., 2024; Hurtado-Mazeyra et al., 2024; Pinhei ro-Carozzo & Murta, 2023; Pontes & Brino, 2022). Second, model adaptation and ensembles: evaluate domain adaptation/fine-tuning on high-quality, human-coded corpora and multi-pass/aggregator prompts that emu late group consensus to capture “social wisdom,” im proving alignment while preserving precision. Third, explain divergence: design studies that trace where expert and AI cues part company (e.g., temporal con tingency, turn-taking repairs, prosodic warmth, distri butional use of anchors) so that alignment can improve without sacrificing internal consistency. Pairwise ICCs (with k = 2 raters per comparison) are inherently less stable than multi-rater estimates and are sensitive to differences in scale variance. By defi ning targets at the video × item level, we maximised item-specific comparability but introduced heteroge neity in pairwise coverage (n), which we report along side each comparison. Bootstrap CIs were therefore preferred to closed-form intervals. Given the modest sample of videos (14), results should be interpreted as a methodological proof-of-concept and revisited with larger, more diverse corpora.
Conclusion Multimodal GenAI (Gemini 2.5) achieves a narrow, highly stable rating profile (high precision) scaling past some hu man evaluators while acting as a consistent yet different evaluator relative to human experts, acting as an in creasingly powerful yet distinctly biased collaborator. In LMIC contexts, this profile is valuable for scale and cost provided AI is embedded in human-in-the-loop (HITL) workflows with cultural calibration, explicit ver sioning, and ongoing validation, aligning with regional evidence on home stimulation, shared reading, and ob servational assessment of interaction quality (Balog et al., 2024; Orozco Restrepo et al., 2022; Pinheiro-Carozzo & Murta, 2023). Equally, process matters: student con sensus outperformed single judges in precision, sug gesting next-generation AI raters should be trained not only on expert labels but also on principles of reliable group judgment. Responsible, human-centred integra tion-pairing efficiency with equity, interpretability, and oversight can broaden access to high quality develop mental assessment, provided that systematic biases are calibrated.
Data and code availability Because recordings involve minors and sensitive perso nal data, raw videos cannot be shared publicly. Regar ding reproducibility, the components of the pipeline differ in their openness: (a) the statistical analyses, in cluding all classical interrater indices and the hierar chical Bayesian model (PyMC), are fully reproducible via the companion Colab notebook (link below), which contains code, synthetic data structures, and sensi tivity analyses without any identifiable media; (b) the preprocessing pipeline—including ASR transcription and computer-vision feature extraction—relies on open-source libraries whose configuration is described in the Method section, but the specific intermediate ou tputs (segment-level transcripts and feature vectors) cannot be shared due to the data-use agreement; and (c) Gemini model access is subject to Google’s API terms, meaning exact outputs may vary across API versions or access tiers. De-identified quantitative data (aggregated rater scores) and full analysis code are available from the corresponding author upon reasonable request and a data-use agreement consistent with ethics approval. A reproducible notebook with the Bayesian model and sensitivity analyses (no identifiable media) is available at:
https://colab.research.google.com/drive/13D8Hco JU0AzIRoit8 uHK_o0JHEOxeGIT?usp=sharing . Conflicts of interest The authors declare no conflicts of interest.
9/11 Evaluating early development with AI: A multimodal analysis of caregiver–toddler interactions in Argentina Funding This study was supported by the Universidad de La Sabana (Grant No. GL:GL0000000001695).
References Anil, R., Borgeaud, S., Alayrac, J.-B., Yu, J., Soricut, R., Schalk wyk, J., Dai, A. M., Hauth, A., Millican, K., Silver, D., John son, M., Antonoglou, I., Schrittwieser, J., Glaese, A., Chen, J., Pitler, E., Lillicrap, T., Lazaridou, A., Firat, O., … & Vinyals, O. (2024). Gemini: A family of highly capable multimodal models. https://arxiv.org/abs/2312.11805 Baevski, A., Zhou, Y., Mohamed, A., & Auli, M. (2020). wav2vec 2.0: A framework for self-supervised learning of speech representations. Advances in Neural Information Processing Systems, 33, 12449-12460. https://proceedings.neurips.cc/pa per/2020/file/92d1e1eb1cd6f9fba3227870bb6d7f07-Paper.pdf Balog, L. G. C., Benitez, P., Costa, A. R. A. da ., & Domeniconi, C. (2024). Shared reading: Interaction between parents and children during storybook reading. Estudos de Psicologia (Campinas), 41, Article e210015. https://doi.org/10.1590/1982- 0275202441e210015 Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., ... & Amodei, D. (2020). Language models are few-shot learners. In H. Larochelle, M. Ranzato, R. Hadsell, M. F. Balcan, & H. Lin (Eds.), Advances in neural information processing sys tems (Vol. 33, pp. 1877–1901). Curran Associates, Inc. https:// proceedings.neurips.cc/paper/2020/file/1457c0d6bfcb 4967418bfb8ac142f64a-Paper.pdf Bruner, J. S. (1978). The role of dialogue in language acquisition. In A. Sinclair, R. Jarvella, & W. J. M. Levelt (Eds.), The child’s conception of language (pp. 241-256). Springer. Carballo-Fazanes, A., Rey, E., Valentini, N. C., Rodríguez-Fernán dez, J. E., Varela-Casal, C., Rico-Díaz, J., Barcala-Furelos, R., & Abelairas-Gómez, C. (2021). Intra-rater and inter-rater (expert vs. novice) reliability of the test of gross motor development— third edition. International Journal of Envi ronmental Research and Public Health, 18(4), 1652. https://doi. org/10.3390/ijerph18041652 Carcone, A. I., Hasan, M., Alexander, G. L., Dong, M., Eggly, S., Brogan Hartlieb, k., Naar, S., MacDonell, K., Kotov, A. (2019). Developing machine learning models for behavioral cod ing. Journal of Pediatric Psychology, 44(3), 289-299. https://do i.org/10.1093/jpepsy/jsy113 CEPAL. (2024, April). Pobreza en américa latina alcanzó el nivel más bajo desde que se tiene registro comparable: es esencial el fortalecimiento de la protección social no contributiva para avanzar hacia el desarrollo social inclusivo [Comunicado de prensa]. https://www.cepal.org /es/comunicados/cepal-latasa-pobreza-regional-que-aumento-la-pandemia-se-hareducid o-un-nivel-similar Chu, S. Y., Kim, J. W., & Yi, M. Y. (2025). Think together and work better: Combining humans’ and LLMs’ think-aloud out comes for effective text evaluation. In Proceedings of the
2025 CHI Conference on Human Factors in Computing Systems
Association for Computing Machinery, 1089, 1-23. https://doi. org/10.1145/3706598.3713181 Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for lan guage understanding. In J. Burstein, C. Doran, & T. Solorio (Eds.), Proceedings of the 2019 Conference of the North Ameri can Chapter of the Association for Computational Linguistics: Human Language Technologies (Vol. 1, pp. 4171–4186). Associ ation for Computational Linguistics. https://aclanthology. org/N19-1423/ Dickinson, D. K., & Porche, M. V. (2011). Relationship between language experiences in preschool classrooms and chil dren’s kindergarten and fourth-grade language and read ing abilities. Child Development, 82(3), 870-886. https://doi. org/10.1111/j.1467-8624.2011.01576.x Giraldo-Huertas, J. J., Sánchez, D. C., & Gutiérrez-Romero, M. F. (2023). Efectos en el desarrollo cognitivo de niños y niñas en condición de riesgo y pobreza multidimensional de dos intervenciones con cuidadores principales. Revista Com plutense de Educación, 34(1), 157-166. https://doi.org/10.5209/ rced.77229 Goertzen, B. J., & Klaus, K. (2023). Is it actually reliable? Ex amining statistical methods for inter-rater reliability of a rubric in graduate education. Research & Practice in As sessment, 18(2), 31-41. https://www.rpajournal.com/is-it-ac tually-reliable-examining-statistical-methods-for-in ter-rater-reliability-of-a-rubric-in-graduate-education/ Goodfellow, I. J., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Far ley, D., Ozair, S., Courville, A., & Bengio, Y. (2014). Generative adversarial nets. In Z. Ghahramani, M. Welling, C. Cortes, N. Lawrence, & K. Q. Weinberger (Eds.), Advances in neural in formation processing systems (Vol. 27, pp. 2672–2680). Curran Associates, Inc. https://papers.nips.cc/paper/5423-genera tive-adversarial-nets Guest, G., Bunce, A., & Johnson, L. (2006). How many inter views are enough?: An experiment with data saturation and variability. Field Methods, 18(1), 59-82. https://doi. org/10.1177/1525822X05279903 Hallgren, K. A. (2012). Computing inter-rater reliability for observational data: An overview and tutorial. Tutorials in Quantitative Methods for Psychology, 8(1), 23-34. https://doi. org/10.20982/tqmp.08.1.p023 Hashemi, H., Eisner, J., Rosset, C., Van Durme, B., & Kedzie, C. (2024). LLM-Rubric: A multidimensional, calibrated ap proach to automated evaluation of natural language texts. In L.-W. Ku, A. Martins, & V. Srikumar (Eds.), Proceedings of the 62nd Annual Meeting of the Association for Computa tional Linguistics (Volume 1: Long Papers; pp. 13806-13834). Association for Computational Linguistics. https://doi. org/10.18653/v1/2024.acl-long.745 Hertzog, M. A. (2008). Considerations in determining sample size for pilot studies. Research in Nursing & Health, 31(2), 180-
191. https://doi.org/10.1002/nur.20247
Hu, Z., Wang, L., Lan, Y., Xu, W., Lim, E.-P., Bing, L., Xu, X., Poria, S., & Lee, R. (2023). LLM-Adapters: An adapter family for pa rameter-efficient fine-tuning of large language models. In H. Bouamor, J. Pino, & K. Bali (Eds.), Proceedings of the 2023 Conference on Empirical Methods in Natural Language Process ing (pp. 5254–5276). Association for Computational Linguis tics. https://doi.org/10.18653/v1/2023.emnlp-main.319 Hurtado-Mazeyra, A., Alejandro-Oviedo, O. M., Ollachica-Hum piri, R., & Borda-Acero, C. (2024). El juego padres-hijo y su relación con el desarrollo cognitivo y socioemocional. Re vista Latinoamericana de Ciencias Sociales, Niñez y Juventud, 22(1), 1-19. https://doi.org/10.11600/rlcsnj.22.1.5875 Idrovo, A. J. (2024). Successful local science in low-income and middle-income countries. Lancet (London, England), 403(10427), 615. https://doi.org/rdgf Inter-American Development Bank (IDB). (2020). Equitable education in Latin America: Policy brief Johnson, K., & Stead, W. (2022). Making electronic health re cords both safer and smarter. JAMA. 328(6), 523-524. https:// doi.org/10.1001/jama.2022.12243 Kollias, D., & Zafeiriou, S. (2021). Analysing affective behav ior in the second ABAW2 competition. In Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops (pp.
3652–3660).
https://doi.org/10.1109/IC
CVW54120.2021.00408
10/11 J. Amorocho et al.
Koopowitz, S.-M., Maré, K. T., Lake, M., du Plooy, C., Hoffman, N., Donald, K. A., Malcolm-Smith, S., Murray, L., Zar, H. J., Coop er, P. J., & Stein, D. J. (2024). Efficacy of a dialogic book-shar ing intervention in a South African birth cohort: A ran domized controlled trial. Comprehensive Psychiatry, 128,
152436. https://doi.org/10.1016/j.comppsych.2023.152436
Lin, Y., Li, R., Ribosa, J., Duran, D., & Sun, B. (2024). Expert and novice teachers’ cognitive neural differences in under standing students’ classroom action intentions. Brain Sci ences, 14(11), 1080. https://doi.org/10.3390/brainsci14111080 Marko, J. G. O., Neagu, C. D., & Anand, P. B. (2025). Examining inclusivity: The use of ai and diverse populations in health and social care: A systematic review. BMC Medical Infor matics and Decision Making, 25(1), 57. https://doi.org/10.1186/ s12911-025-02884-1 Mol, S. E., & Bus, A. G. (2011). To read or not to read: A meta-anal ysis of print exposure from infancy to early adulthood. Psychological Bulletin, 137(2), 267-296. https://doi.org/10.1037/ a0021890 Montirosso, R., Castagna, A., Butti, N., Innocenti, M. S., Rogg man, L. A., & Rosa, E. (2023). A contribution to the Italian validation of the parenting interaction with children: Checklist of observations linked to outcome (piccolo). Frontiers in Psychology, 14, 1105218. https://doi.org/10.3389/ fpsyg.2023.1105218 Motiwalla, L. F. (2007). Mobile learning: A framework and eval uation. Computers & Education, 49(3), 581-596. https://doi. org/10.1016/j.compedu.2005.10.011 Murray, L., Jennings, S., Perry, H., Andrews, M., De Wilde, K., Newell, A., Mortimer, A., Phillips, E., Liu, X., Hughes, C., Melhuish, E., De Pascalis, L., Dishington, C., Duncan, J., & Cooper, P. J. (2023). Effects of training parents in dialog ic book-sharing: The early-years provision in children’s centers (EPICC) study. Early Childhood Research Quarterly, 62, 1-16. https://doi.org/10.1016/j.ecresq.2022.07.008 Murray, L., Rayson, H., Ferrari, P.-F., Wass, S. V., & Cooper, P. J. (2022). Dialogic book-sharing as a privileged intersubjec tive space. Frontiers in Psychology, 13, 786991. https://doi. org/10.3389/fpsyg.2022.786991 Nagrani, A., Yang, S., Arnab, A., Jansen, A., Schmid, C., & Sun, C. (2021). Attention bottlenecks for multimodal fusion. Advances in Neural Information Processing Systems 34 (NeurIPS 2021). https://proceedings.neurips.cc/paper/2021/ hash/76ba9f564ebbc35b1014ac498fafadd0-Abstract.html Norrie, C. S., Deckers, S. R. J. M., Radstaake, M., & van Balkom, H. (2024). A narrative review of the sociotechnical landscape and potential of computer-assisted dynamic assessment for children with communication support needs. Mul timodal Technologies and Interaction, 8(5), 38. https://doi. org/10.3390/mti8050038 Ong, J. C., Ning, Y., Liu, M., Ma, Y., Liang, Z., Singh, K., Chang, R. T., Vogel, S., Lim, J. C., Tan, I. S., Freyer, O., Gilbert, S., Bit terman, D. S., Liu, X., Denniston, A. K., & Liu, N. (2025). Reg ulatory science innovation for generative AI and large language models in health and medicine: A global call for action. ArXiv. https://arxiv.org/abs/2502.07794 Oremus, M., Oremus, C., Hall, G. B. C., & McKinnon, M. C. (2012). Inter-rater and test–retest reliability of quality assess ments by novice student raters using the Jadad and New castle–Ottawa scales. BMJ Open, 2(4), Article e001368. https://doi.org/10.1136/bmjopen-2012-001368 Orozco Restrepo, L. A., Cardona Cañas, M. F., & Barrios Arro yave, F. A. (2022). Estimulación temprana en el hogar de infantes que asisten a un centro infantil. Revista Cuidarte, 13(1), Article e2142. https://doi.org/10.15649/cuidarte.2142 Pathak, A., Gandhi, R., Uttam, V., Ramamoorthy, A., Ghosh, P., Jindal, A. R., Verma, S., Mittal, A., Ased, A., Khatri, C., Nakka, Y., Devansh, Challa, J. S., & Kumar, D. (2025). Rubric is all you need: Improving LLM-based code evaluation with ques tion-specific rubrics. In Proceedings of the 2025 ACM Confer ence on International Computing Education Research V.1 (pp. 181–195). Association for Computing Machinery. https://doi. org/10.1145/3702652.3744220 Pelau, C., Dabija, D.-C., & Ene, I. (2021). What makes an AI device human-like? the role of interaction quality, empathy and perceived psychological anthropomorphic characteristics in the acceptance of artificial intelligence in the service industry. Computers in Human Behavior, 122, 106855. https:// doi.org/10.1016/j.chb.2021.106855 Pike, W. A. (2005). Augmenting collaboration through situated representations of scientific knowledge. [Tesis Doctoral, The Pennsylvania State University]. https://goo.su/SNHTQ9 Pinheiro-Carozzo, N. P., Murta, S. G., & Souza, A. S. S. (2023). Application of direct and systematic observation of in teraction with mother-adolescent child dyads. Paidéia (Ri beirão Preto), 33, Article e3323. https://doi.org/10.1590/1982- 4327e3323 Pontes, V. I. P., & Brino, A. L. F. (2022). A pilot study with obser vational measures of parent–child interaction in clini cal context. Psicologia: Teoria e Pesquisa, 38, Article e38313. https://doi.org/10.1590/0102.3772e38313.en Radanliev, P. (2024). Artificial intelligence: Reflecting on the past and looking towards the next paradigm shift. Jour nal of Experimental & Theoretical Artificial Intelligence, 37(7), 1045–1062. https://doi.org/10.1080/0952813X.2024.2323042 Radford, A., Kim, J. W., Xu, T., Brockman, G., McLeavey, C., & Sut skever, I. (2023). Robust speech recognition via large-scale weak supervision. Proceedings of the 40th International Con ference on Machine Learning, 202, 28492-28518. https://pro ceedings.mlr.press/v202/radford23a.html Radford, A., Narasimhan, K., Salimans, T., Sutskever, I. (2018). Improving language understanding by generative pre-training. https://www.mikecaptain.com/resources/pdf/GPT-1.pdf Saab, K., Tu, T., Weng, W.-H., Tanno, R., Stutz, D., Wulczyn, E., Zhang, F., Strother, T., Park, C., Vedadi, E., Chaves, J. Z., Hu, S.-Y., Schaekermann, M., Kamath, A., Cheng, Y., Barrett, D. G. T., Cheung, C., Mustafa, B., Palepu, A., …& Natarajan, V. (2024). Capabilities of Gemini models in medicine. https://arx iv.org/abs/2404.18416 Sahdra, B. K., King, G., Payne, J. S., Ruiz, F. J., Ali Kolahdou zan, S., Ciarrochi, J., & Hayes, S. C. (2024). Why research from lower- and middle-income countries matters to ev idence-based intervention: A state of the science review of act research as an example. Behavior Therapy, 55(6), 1348-
1363. https://doi.org/10.1016/j.beth.2024.06.003
Santelices, M. P., Carvacho, C., Farkas, C., León, F., Galleguillos, F., & Himmel, E. (2012). Medición de la sensibilidad del adul to con niños de 6 a 36 meses de edad: construcción y análi sis preliminares de la escala de sensibilidad del adulto, esa. Terapia Psicológica, 30(3), 19-29. https://doi.org/10.4067/ S0718-48082012000300003 Shafiee Rad, H. (2025). Reinforcing L2 reading comprehension through artificial intelligence intervention: Refining en gagement to foster self-regulated learning. Smart Learning Environments, 12(1), 23. https://doi.org/10.1186/s40561-025-
00377-2
Su, J., & Yang, W. (2022). Artificial intelligence in early child hood education: A scoping review. Computers and Education. Artificial Intelligence, 3, 100049. https://doi.org/10.1016/j. caeai.2022.100049 Temple, J. (2024). Using AI for qualitative labeling: Consistency and comparisons [Honors Program Theses. 240]. https://schol arship.rollins.edu/honors/240
11/11 Evaluating early development with AI: A multimodal analysis of caregiver–toddler interactions in Argentina Vally, Z., Murray, L., Tomlinson, M., & Cooper, P. J. (2015). The impact of dialogic book-sharing training on infant lan guage and attention: A randomized controlled trial in a deprived South African community. Journal of Child Psy chology and Psychiatry, 56(8), 865-873. https://doi.org/10.1111/ jcpp.12352 Vally, Z. (2012). Dialogic reading and child language growth— combating developmental risk in South Africa. South African Journal of Psychology, 42(4), 617-627. https://doi. org/10.1177/008124631204200415 Van Noorden, R., & Perkel, J. M. (2023). AI and science: What 1,600 researchers think. Nature, 621, 672-675. https://doi. org/10.1038/d41586-023-02980-0 Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., & Polosukhin, I. (2017). Attention is all you need. NeurIPS 2017. https://papers.nips.cc/pa per/7181-attention-is-all-you-need.pdf Vilaseca, R., Rivero, M., Bersabé, R. M., Navarro-Pardo, E., Can tero, M. J., Ferrer, F., Valls Vidal, C., Innocenti, M. S., & Rog gman, L. (2019). The Spanish version of the piccolo (parent ing interactions with children: Checklist of observations linked to outcomes): A validation study. Frontiers in Psy chology, 10, 680. https://doi.org/10.3389/fpsyg.2019.00680 Vygotsky, L. S. (1978). Mind in society: The development of higher psychological processes. Harvard University Press. Wang, S., Sun, Z., Li, M., Zhang, H., & Metwally, A. H. S. (2024). Leveraging tik-tok for active learning in management ed ucation: An extended technology acceptance model ap proach. The International Journal of Management Education, 22(3), 101009. https://doi.org/10.1016/j.ijme.2024.101009 Wind, S. A. (2019). Examining the impacts of rater effects in performance assessments. Applied Psychological Measure ment, 43(2), 159-171. https://doi.org/10.1177/0146621618789391 World Bank. (2021). Education and literacy in Latin America: Re gional statistics and insights.
Wu, X., Saraf, P. P., Lee, G., Latif, E., Liu, N., & Zhai, X. (2025). Unveiling scoring processes: Dissecting the differences between llms and human graders in automatic scoring. Technology, Knowledge and Learning, 31, 669-684. https://doi. org/10.1007/s10758-025-09836-8 Xames, M. D. (2025). Generative artificial intelligence is emerg ing as a phantom doctor in low- and middle-income coun tries - health systems must respond. Journal of Medical Sys tems, 49(1), 86. https://doi.org/10.1007/s10916-025-02223-x
Cita: Amorocho, José, Giraldo-Huertas, Juan José, Gago-Galvagno, Lucas G., Elgier, Angel M., Mancini, Natalia A. (2026), Evaluating early development with AI: A multimodal analysis of caregiver–toddler interactions in Argentina, Fundación Universitaria Konrad Lorenz, p. N. https://repositorio.konradlorenz.edu.co/handle/001/7833