Skip to main content

Association for Vietnamese Language and Speech Processing

A chapter of VAIP - Vietnam Association for Information Processing

Numerical Reasoning QA

Important Dates

September 18: Training Data Release

September 25: Public Test Release

October 9: Private Test Release

October 15: Result Announcement

October 25: Paper Submission

November 5: Acceptance Notification

November 12: Camera-ready Deadline

November 15: Workshop Date

Task Description

This shared task focuses on multimodal numerical reasoning over Vietnamese financial documents. Participants are required to develop systems capable of understanding and integrating information from multiple modalities, including textual content, tables, and chart images, in order to answer numerical questions about financial reports.

Unlike conventional financial question answering tasks that mainly rely on textual and tabular information, this task extends numerical reasoning to multimodal financial documents. Relevant evidence may appear in text, tables, charts, or across multiple modalities, requiring systems to identify the appropriate information and perform the necessary reasoning operations.

The expected output of a system is a reasoning program representing the sequence of operations required to solve the given question. This program-based representation makes the reasoning process explicit and enables the correctness of the reasoning procedure to be evaluated directly.

Through this shared task, we aim to promote research on multimodal financial document understanding, numerical reasoning, and interpretable reasoning systems for Vietnamese.

Dataset

The raw data is collected from publicly available financial reports of Vietnamese enterprises during the 2020–2025 period.

The source documents include annual reports, industry analysis reports, company analysis reports, and strategic reports containing financial information presented in multiple forms.

The dataset integrates three main information modalities:

  • Text: textual passages containing descriptions, explanations, and financial information.
  • Table: structured tabular data containing financial indicators and numerical values.
  • Image: chart images and other visual representations of financial information.

Each question may require evidence from a single modality or the integration of information across multiple modalities. In particular, some questions require numerical information available in charts together with contextual information provided by text or tables.

Dataset Format

The dataset is provided in JSON format. Each sample contains four main fields:

{
  "text": textual context associated with the question,
  "table": tabular information associated with the document,
  "image": chart image or visual information associated with the document,
  "question": the natural-language numerical reasoning question
}

An example of the data structure is shown below:

{
  "text": [...],
  "table": [...],
  "image": "...",
  "question": "..."
}

The text, table, and image fields provide the multimodal context from which the system must retrieve relevant numerical evidence. Given this context and the corresponding question, participants are required to generate the reasoning program necessary to solve the problem.

Reference reasoning programs for the evaluation sets are retained by the organizers and are used only for automatic evaluation.

Evaluation Metric

Systems are evaluated using Program Accuracy.

Program Accuracy measures whether the reasoning program generated by a system correctly represents the reasoning procedure required to answer the question. The predicted program is compared against the reference reasoning program provided by the organizers using the official evaluation script.

This metric evaluates the reasoning process itself rather than considering only the final numerical answer. A system therefore needs to identify the correct numerical evidence, select the appropriate reasoning operators, and construct the correct sequence of operations.

The final ranking of participating systems is determined based on their Program Accuracy on the private test set.

Submission Format

Participants are required to submit the predicted reasoning program for each sample in the evaluation set.

Each prediction must correspond to one input question and follow the reasoning-program syntax defined by the shared task.

The detailed JSON submission format, supported operators, program grammar, and submission instructions will be provided together with the dataset and official evaluation script.

Data Usage

Participants may use the provided training data to develop their systems. Additional publicly available or appropriately licensed resources may also be used, provided that their use is clearly described in the submitted system paper.

Task setting and participation rules

The Multimodal Reasoning challenge consists of a single official subtask. The challenge is designed to encourage the development of efficient multimodal reasoning systems using small and medium-sized models.

  • Single subtask: The challenge contains one official subtask.
  • Model size constraint: Each individual model used in the submitted system must have no more than 8 billion parameters (≤ 8B). Systems may consist of multiple model components or multiple stages, provided that no individual model exceeds the 8B parameter limit. For example, a pipeline consisting of two 7B models is permitted.
  • External data: The use of additional external data is permitted. Teams may use publicly available or self-constructed datasets in addition to the data released by the organizers.
  • External data disclosure: All external datasets and additional data resources used for training, fine-tuning, data augmentation, retrieval, or other stages of the submitted system must be clearly documented in the team's technical report. Teams must provide the source of the data, a brief description of the data, and an explanation of how the data were used.
  • External data submission: Any external data used by the system must also be submitted or made accessible to the organizers together with the technical report as supplementary material. For publicly available datasets, teams may provide the corresponding dataset references and access links. For self-constructed or modified datasets, teams must provide the data or sufficient materials for the organizers to verify its use.
  • Official ranking: Compliance with both the model-size constraint and the external-data disclosure requirements is mandatory for inclusion in the official leaderboard and final ranking. Teams that do not provide sufficient information about the models or external data used in their submitted systems may be excluded from the official ranking.
Contact

Zalo Group: []

Registration

https://forms.gle/RTU4ncNog5nmis28A 

Organizers
  • Nguyen Thi Minh Huyen
  • Ha My Linh
  • Le Ngoc Toan
  • Dang Phuong Nam
  • Vu Xuan Luong
  • Pham Thi Duc
  • Ngo The Quyen
  • Phan Thi Hue
  • Le Van Cuong
References
  1. Aida Amini, Saadia Gabriel, Shanchuan Lin, Rik Koncel-Kedziorski, Yejin Choi, and Hannaneh Hajishirzi. 2019. MathQA: Towards interpretable math word problem solving with operation-based formalisms. In Proceedings of NAACL-HLT 2019.
  2. Zhiyu Chen et al. 2021. FinQA: A Dataset of Numerical Reasoning over Financial Data. In Proceedings of EMNLP 2021.
  3. Ahmed Masry, Do Xuan Long, Jia Qing Tan, Shafiq Joty, and Enamul Hoque. 2022. ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning. In Findings of ACL 2022.

Sponsors and Partners

VinBIGDATA   VinIF  AIMESOFT  bee  Dagoras            

 

  zalo    VTCC  VCCorp

 

 

IOIT  HUS  USTH  UET    TLU  UIT  INT2  jaist  VIETLEX