Skip to main content

Association for Vietnamese Language and Speech Processing

A chapter of VAIP - Vietnam Association for Information Processing

Speaker Recognition on Edge device

Important dates

September 18: Training Data release

September 25: Public Test

October 9: Private Test

October 15: Result announcement

October 25: Paper submission

November 5: Acceptance notification

November 12: Camera-ready (= > Proceedings)

November 15: Workshop date

General Description

Speaker Recognition (SR) is the task of identifying or verifying a speaker's identity based on voice characteristics. It is typically framed in two forms: speaker identification (determining which speaker, among a known set, is speaking) and speaker verification (determining whether two audio segments belong to the same speaker). This challenge focuses on the speaker verification setting: given a pair of audio files, the system must output a score reflecting the similarity between the two voices, from which a same-speaker / different-speaker decision can be derived.

Modern SR models (ResNet-based architectures, ECAPA-TDNN, self-supervised backbones such as wav2vec2/XLS-R, etc.) achieve strong accuracy on clean, curated data. However, when deployed on edge devices (smartphones, IoT devices, embedded chips with limited compute), systems face several challenges that idealized training conditions often fail to capture:

  • Limited computational and memory resources: large, high-parameter-count models are difficult to run in real time on edge CPUs/NPUs with constrained RAM and strict power-consumption budgets.
  • Degraded audio quality in real-world conditions: low-cost microphones, background noise, reverberation, variable speaking distance, and codec compression during transmission.
  • Diverse recording conditions: from professional studio recordings to casually captured real-world audio, requiring models with strong generalization rather than models overfit to a single domain.
  • Spoofing risk: in some extended scenarios, edge SR systems must also consider robustness against spoofing attacks (replay, synthesis, voice conversion). This aspect is optional in this challenge and not a mandatory requirement for all teams.

The goal of this challenge is to encourage participants to build speaker verification systems that are both accurate and resource-efficient, capable of operating reliably under realistic acoustic conditions while remaining compatible with edge-device deployment constraints.

Dataset

The dataset is designed to closely reflect real-world edge-device deployment conditions and consists of two main components:

  1. Main data: collected from social media sources, reflecting diverse real-world recording conditions (background noise, inconsistent microphone quality, varied recording contexts).
  2. Edge-device subset: a portion of the data is recorded specifically using edge devices, more accurately simulating the acoustic characteristics of audio captured directly on constrained hardware (e.g., embedded-device microphone quality, speaking distance, real deployment environments).

A labeled training set will be provided in full from the start of the competition for model development. The public test set and private test set will be released progressively through the competition platform; participants are responsible for monitoring announcements to stay up to date with release schedules.

Evaluation Metrics

Systems will be evaluated using Equal Error Rate (EER) — a standard metric for speaker verification, representing the operating threshold at which the False Acceptance Rate equals the False Rejection Rate. A lower EER indicates better performance.

Submission Format

Participants must submit a plain-text result file (e.g., .txt), with each line corresponding to one audio pair to be compared, in the following format:

file1.wav file2.wav score

Where:

  • file1.wav, file2.wav: the file names (or relative paths) of the two audio files being compared, matching exactly the pair list provided in the test trial file.
  • score: a real-valued number representing the similarity between the two voices (e.g., cosine similarity or a log-likelihood ratio, depending on the participant's system — no fixed range is required). A higher score indicates a higher likelihood that the two segments belong to the same speaker.

Each line should be separated by a newline, with fields within a line separated by a single space. Submissions must include all trial pairs, in the same order as provided in the corresponding test trial list, to avoid errors during automated scoring.

Registration

Participants can register through this link: https://forms.gle/fD1YaaPkoHKdsyZA8

Organizer

Sponsors and Partners

VinBIGDATA   VinIF  AIMESOFT  bee  Dagoras            

 

  zalo    VTCC  VCCorp

 

 

IOIT  HUS  USTH  UET    TLU  UIT  INT2  jaist  VIETLEX