Sapia.ai compared: text-first chat screening versus a conversational AI panel
Sapia.ai made a deliberate and defensible design choice: remove video entirely, and with it appearance, background and camera quality as assessment inputs. It is a real fairness advantage, and it comes with a real trade-off in evidence depth.
- Both products remove video from the assessment; they differ in how much evidence the format produces.
- Sapia’s short written answers are mobile-friendly and fast — well suited to high-volume frontline screening.
- A conversation with adaptive probing produces more evidence per competency, which matters as roles get more complex.
- Modality choice — voice, video or text — is an accessibility question as much as a preference.
- For either, the Australian questions are the same: per-decision evidence, APP 1.7 classification and six-year retention.
The difference: how much evidence the format produces
Sapia’s public description of its format is a structured chat interview of roughly five questions, with candidates typically giving five to seven text-based behavioural answers of around 50 to 150 words. That is deliberately light on the candidate and produces consistent, structured, comparable data quickly.
FirstPanel runs a conversation. Eight agents each own a competency, follow-ups are generated adaptively when an answer is thin, and the candidate can respond in voice, video or text in whichever of 40+ languages they are strongest in.
| Short-form text screening | Conversational AI panel | |
|---|---|---|
| Candidate effort | Low — a few short written answers on a phone | Moderate — a genuine interview, on their own schedule |
| Evidence per competency | Bounded by the written answer length | Extended by adaptive probing until there is enough to rate or an abstention is recorded |
| Thin answers | Scored as given | Probed with a follow-up before scoring |
| Modality | Text | Voice, video or text, candidate’s choice |
| Written-fluency dependence | Higher — the assessment is entirely written | Lower — a candidate who speaks better than they type can choose voice |
| Best fit | Very high-volume frontline screening, mobile-first | Roles where depth, probing and evidence density matter |
Neither column is simply better. A short written screen that four hundred people complete on their phones in six minutes is a genuinely strong instrument for the top of a frontline funnel. It is a weaker one for a role where the competency you most need to assess only shows up in the third follow-up.
The written-fluency question
A text-only assessment removes appearance and accent, which is the point. It substitutes written expression, which is not neutral either — it correlates with education, with first-language status, and with disabilities affecting writing.
That is a trade, not a flaw, and it is a favourable trade for many roles. It is worth being explicit about, because "we removed bias by removing video" is a claim about one input, not about the whole assessment.
- For roles where written communication is a genuine job requirement, assessing it is job-related and appropriate.
- For roles where the job is spoken — floor service, care work, cabin crew, contact centre — a written-only assessment measures something adjacent to the job rather than the job.
- Letting the candidate choose modality removes the trade rather than swapping which group it disadvantages, which is why we made modality a candidate choice rather than a product decision.
- Whichever you use, monitor adverse impact per requisition — the format’s intended fairness property is a hypothesis about your pipeline until you have measured it on your pipeline.
Choosing between them
Where we would point you at Sapia.ai rather than at us:
- Extremely high-volume frontline screening where the binding constraint is candidate completion on a phone in under ten minutes.
- A pipeline where written English is a genuine and stated job requirement.
- A buyer who wants the longest available Australian track record in text-first screening specifically.
Where we would argue for FirstPanel:
- Roles where the predictive competencies need probing to surface — judgement, composure, integrity — rather than being visible in a short written answer.
- Multilingual or spoken-role pipelines where written fluency is not the job.
- Where the per-decision record has to carry a Fair Work reverse-onus defence: verbatim evidence per rating, explicit abstentions, rubric and model versioning, named human decision-maker, six-year retention.
- Seasonal demand where per-completed-interview pricing beats a licence you carry between peaks.