How Kenyan students actually experience Gemini as a study companion — and where it still falls short of feeling local.
AI tools are quickly becoming part of how students study, plan, and cope with the pressures of academic life. But most conversations about AI in education still centre on a narrow question: is the answer correct? Far less attention is paid to whether that answer actually feels useful, trustworthy, and relevant to the student receiving it.
At YUX, we set out to explore this gap through a three-day Live LLM Testing study of Google's Gemini-2.0-flash, run with students in Kenya through Kitala.ai, our in-house AI evaluation platform. Rather than asking whether Gemini could answer academic questions correctly, we wanted to understand how it performed across the full, messy, multi-turn conversations that make up a student's real academic and personal life — from career doubts and side hustles to declining grades and academic pressure.
This mattered to us because a large share of AI evaluation today still relies on benchmark performance and one-off question-answer accuracy. These methods say little about how a model handles a student who is anxious, code-switching between English, Swahili, and Sheng, or returning to the same conversation over several turns. We wanted to know: does Gemini hold up once the conversation gets complicated?

Participants seated in a discussion circle
Methodology
The study combined live, multi-turn conversations with structured human evaluation, designed to mirror how students actually use AI rather than how they are typically tested on it.
Live LLM testing conducted over 3 days
8 education scenarios, developed through a design workshop with education and youth experts, grouped under three themes: Looking to the Future, Study & Assignments, and Day in the Life of a Student
60 human-based evaluations, scored across four metrics: Accuracy, Clarity, Cultural Relevance, and Trustworthiness
20 participants, split between AI Power Users (regular, confident AI users) and AI Regular Users (occasional or first-time users)
- Model tested: Gemini-2.0-flash, no system prompt, temperature 0.8

Walkthrough of the Eight Education Scenarios. Participants were introduced to the eight real-world education scenarios that formed the basis of the evaluation.

A Kitala.ai session showing a live, multi-turn conversation on career uncertainty, evaluated turn-by-turn across all four metrics.
Overall Score
Across all 60 evaluations, Gemini performed consistently well, but one metric stood out as the clear laggard.
Accuracy: 4.2 / 5.0
Clarity: 4.3 / 5.0
Cultural Fit: 4.0 / 5.0
Trustworthiness: 4.2 / 5.0
Cultural relevance was the lowest-scoring dimension overall, echoing a pattern we've observed in previous studies: models can be accurate and clear while still feeling slightly foreign to the person on the other end of the conversation. Segmenting the results further, AI Regular Users rated the model marginally higher on average (4.3) than AI Power Users (4.2), suggesting that the more experience a student has with AI tools, the more critically they evaluate it.

Overall scores across all 60 evaluations, and segment rankings by AI Regular Users vs AI Power Users.
Score per Use Case
Breaking the results down by scenario revealed a more nuanced picture than the overall averages suggest. While the model performed consistently well on practical, task-oriented scenarios, it was less consistent in conversations requiring greater contextual and emotional understanding.
Career Uncertainty scored lowest overall, with Cultural Relevance and Trustworthiness both scoring 3.6.
Side Hustle Dilemma followed a similar pattern, with Cultural Relevance scoring 3.2 despite strong Accuracy and Clarity scores (4.5 each), suggesting the advice did not always reflect the student's context.
Repeated Low Marks, a scenario involving declining academic performance, also scored 3.6 for Cultural Relevance, despite achieving the highest Accuracy score in the study (4.6).
By contrast, practical, forward-looking scenarios like Job Application Fatigue and Academic Pressure scored strongly across every metric, including Cultural Relevance (4.6 and 4.0 respectively)

Scores across all 8 education scenarios, grouped by theme.
The pattern is clear: Gemini's weakest moments were not about getting facts wrong. They happened in conversations that required emotional nuance, local slang, or an understanding of what it actually means to be a Kenyan student weighing career choices, family expectations, and financial pressure at once.
Segments vs Use Cases
Looking at how AI Regular Users and AI Power Users scored the same scenarios differently surfaced one of the study's most interesting findings.

Segment-level scores across scenarios. Citation Support and Repeated Low Marks show the widest gaps between AI Regular Users and AI Power Users.
On Citation Support, AI Power Users rated the model notably higher (4.3) than AI Regular Users (3.8) — likely because power users have a clearer sense of what a strong research response should include, and could better appreciate what Gemini got right. Repeated Low Marks showed the opposite pattern: AI Regular Users rated it far more generously (4.9) than Power Users (4.0), suggesting that less experienced users were more forgiving of generic, reassurance-style advice, while power users expected more tailored guidance for a recurring academic problem.
Qualitative Feedback: What Students Actually Said
Beyond the scores, participants' own words point to two consistent themes: a desire for more warmth, and mixed success with local language.

Participant feedback on empathy and research guidance.

Participant feedback on response length and use of Sheng.

Participant feedback on Kenyan slang and direct responses.
Key Insights
1. Cultural relevance is the quiet weak point
Across nearly every measure, Gemini performed well on accuracy and clarity. But cultural relevance was the one metric that consistently dragged scores down, and it dragged them down hardest in the scenarios that mattered most emotionally: career uncertainty, side hustles, and repeated low marks. Students weren't flagging factual errors; they were flagging a sense that the advice, while correct, didn't quite understand their world.
2. Local language use is inconsistent, not absent
Gemini clearly attempts to use Sheng and Kenyan slang, and when it lands, students notice and appreciate it. But when it misses, the effect is jarring rather than neutral — participants described specific words as “out of context,” which can undermine trust faster than simply avoiding local language altogether. This suggests that attempting cultural adaptation without getting it fully right may carry its own risk.
3. Students want warmth, not just information
The Citation Support feedback captures a recurring tension: participants wanted Gemini to acknowledge confusion or difficulty before jumping into a solution. Fast, information-dense responses were appreciated for efficiency, but students consistently asked for a more human, empathetic tone, particularly in scenarios tied to academic struggle or self-doubt.
4. Response length and format matter as much as content
Several participants, particularly AI Regular Users, pushed back on the length and density of Gemini's responses, asking for information to be “broken down” into a more digestible, eye-appealing format. This was echoed in the Isolation from Friends feedback, where a participant appreciated directness precisely because it meant less reading. Clarity, in other words, isn't only about correctness of language — it's also about how much text a student is willing to read.
Conclusion
These findings highlight that strong technical performance alone is not enough for AI systems to provide meaningful support. While Gemini consistently delivered accurate and clear responses, scenarios involving uncertainty, identity, and personal challenges exposed the importance of cultural relevance, empathetic communication, and contextual understanding. As AI becomes increasingly integrated into students' learning journeys, success will depend not only on producing correct answers but also on delivering responses that resonate with users' lived experiences.
This evaluation was conducted by YUX and Kitala.ai in Kenya in March 2026. Participant details have been anonymised.