Workshop · Deep Learning Indaba 2026 · Lagos
How Do We Know AI Works?
Lessons from Real Evaluations
A working session on evaluating AI systems. You will use a live AI advisor, find where it fails, turn those failures into tests, and fix them in the room.
Organized by APHRC and The Agency Fund.
What we will do
- 01
Break it
Ask the advisor questions from your own work. Read the answer and its sources, then rate it.
- 02
Write the ideal answer
Mark a bad reply and say what it should have said instead. That is the eval case.
- 03
Fix it live
We graduate the best cases into permanent tests, edit the knowledge base on stage, and rerun them.
- 04
Compare two models
Conversations are split 50/50 between two models. Your ratings feed a live A/B test with a real power analysis.
Who it is for
Anyone building, funding, researching, or governing an AI system that people rely on. No evaluation experience needed. Bring one hard question from your own domain.
Facilitators
-
Agnes Kiragga
Speaker and facilitator · APHRC
-
Samuel Iddi
Speaker and facilitator · APHRC
-
Steve Bicko Cygu
Speaker and facilitator · APHRC
-
Edmund Korley
Speaker and facilitator · The Agency Fund
-
Miranda Barasa
Organizer · APHRC
-
Christine Ger Ochola
Organizer · APHRC
What we run on
- Lantern
The advisor you will be testing.
- Marked Problems board
Where the room's findings land.
- Live experiment results
The A/B test, updating as you rate.
- AI Evaluation Playbook
The framework behind the session.
Evaluations run in Calibrate. The experiment runs in Evidential.
Day, time, and room will be confirmed in the Indaba programme. Questions: edmund@agency.fund.