Dataset Viewer
Auto-converted to Parquet Duplicate
Search is not available for this dataset
audio
audioduration (s)
688
1.12k
label
class label
10 classes
0japanese-roleplay-travel-agency01
0japanese-roleplay-travel-agency01
0japanese-roleplay-travel-agency01
1japanese-roleplay-travel-agency02
1japanese-roleplay-travel-agency02
1japanese-roleplay-travel-agency02
2japanese-roleplay-travel-agency03
2japanese-roleplay-travel-agency03
2japanese-roleplay-travel-agency03
3japanese-roleplay-travel-agency04
3japanese-roleplay-travel-agency04
3japanese-roleplay-travel-agency04
4japanese-roleplay-travel-agency05
4japanese-roleplay-travel-agency05
4japanese-roleplay-travel-agency05
5japanese-roleplay-travel-agency06
5japanese-roleplay-travel-agency06
5japanese-roleplay-travel-agency06
6japanese-roleplay-travel-agency07
6japanese-roleplay-travel-agency07
6japanese-roleplay-travel-agency07
7japanese-roleplay-travel-agency08
7japanese-roleplay-travel-agency08
7japanese-roleplay-travel-agency08
8japanese-roleplay-travel-agency09
8japanese-roleplay-travel-agency09
8japanese-roleplay-travel-agency09
9japanese-roleplay-travel-agency10
9japanese-roleplay-travel-agency10
9japanese-roleplay-travel-agency10

Japanese Travel Agency Roleplay Dialogue Corpus

Overview

The Japanese Travel Agency Roleplay Dialogue Corpus is a collection of ten Japanese role-play dialogues simulating travel-agency consultations between a staff member and a customer.

The conversations were collected for speech and dialogue research and include two-speaker mixed audio, speaker-separated audio, and manually created, time-aligned ELAN (.eaf) annotation files.

The dialogues cover a variety of realistic travel consultation scenarios, including destination selection, itinerary planning, transportation, sightseeing, accommodations, and local recommendations. Although designed to resemble natural customer interactions, the recordings were collected through role-playing and are not recordings of actual customer-service conversations.

This repository serves as a public sample demonstrating IngCrowd's expertise in Japanese dialogue-data collection, multi-track audio recording, manual transcription, and ELAN-based time-aligned annotation.


Publicly Available Data

The public dataset contains:

  • 10 Japanese travel-agency consultation dialogues
  • 10 two-speaker mixed audio recordings
  • 20 speaker-separated audio recordings
  • 10 manually created ELAN (.eaf) annotation files
  • Approximately 126 minutes of recorded dialogue

Audio Format

  • WAV
  • 16 kHz sampling rate
  • 16-bit
  • Mono

Video recordings are not included in this repository.

Corresponding video recordings may be provided for research purposes upon request, subject to availability and applicable conditions.


Key Features

  • Japanese role-play dialogues between a travel-agency staff member and a customer
  • Two-speaker mixed audio and speaker-separated audio
  • Manually created ELAN annotations
  • Human-generated verbatim transcriptions
  • Speaker-level time-aligned segmentation
  • Preservation of spontaneous speech phenomena such as fillers, hesitations, repetitions, laughter, and unintelligible speech
  • Diverse travel consultation scenarios covering destinations, transportation, accommodations, sightseeing, and travel planning
  • Suitable for speech processing, dialogue systems, and Japanese language research

Data Collection

The dialogues were collected through role-playing conversations conducted via Zoom.

Each dialogue consists of one participant acting as a travel agency receptionist and another acting as a customer.

Participants joined an online Zoom meeting and were instructed to perform natural conversations based on a predefined travel agency scenario.

Scenario:

  • Customer visits a travel agency.
  • Customer consults about travel plans.
  • Receptionist introduces travel products and provides recommendations.

Dialogue Scenarios

ID Destination / Scenario Staff Customer Duration
01 Okinawa and Naha; attractions, airport access, and rental-car planning Female, 40s Male, 20s 14:07
02 Atami, Shizuoka; a summer family trip with parents and grandparents Female, 40s Male, 20s 12:16
03 Hakone, Kanagawa; family trip with hot springs, accommodations, and sightseeing Female, 40s Female, 40s 12:16
04 Chiba; one-day sightseeing trip for three family members Female, 40s Female, 40s 18:41
05 Ibaraki; spring trip with a grandparent Female, 40s Male, 20s 13:30
06 Niigata; sightseeing consultation for a trip with a parent Female, 40s Male, 20s 12:23
07 Hokkaido and Otaru; trip for three friends Female, 40s Female, 30s 11:31
08 Okinawa and the Yaeyama Islands; summer trip with a parent Female, 40s Female, 30s 11:27
09 Tokyo; planning sightseeing for a friend visiting Japan Female, 40s Female, 30s 13:04
10 Kagoshima; planning a solo spring trip Female, 40s Female, 30s 12:19

Participant information is presented only at a broad demographic level. No personally identifiable information is included.


Data Structure

Each dialogue directory contains the following files:

japanese-roleplay-travel-agencyXX/
├── japanese-roleplay-travel-agencyXX_audio.wav
├── japanese-roleplay-travel-agencyXX_individual_audio1.wav
├── japanese-roleplay-travel-agencyXX_individual_audio2.wav
└── japanese-roleplay-travel-agencyXX.eaf
  • _audio.wav — two-speaker mixed audio
  • _individual_audio1.wav and _individual_audio2.wav — speaker-separated audio
  • .eaf — manually created ELAN annotation with time-aligned transcription

Additional Resources

The dataset includes audio recordings and ELAN annotation files. The corresponding video recordings are not included in this repository. Researchers who wish to obtain the synchronized video recordings for academic or research purposes are encouraged to contact us via email.


Transcription and Annotation

All transcriptions were manually created by human annotators using ELAN.

Each utterance was manually segmented and aligned with the corresponding speech.

Rather than editing spoken language into written language, the annotations preserve fillers, hesitations, repetitions, laughter, and unintelligible speech as faithfully as possible.

No automatic speech recognition (ASR) system was used during transcription or annotation.

The primary annotation tiers are:

  • staff — travel-agency staff member
  • customer — customer seeking travel information

Potential Uses

This dataset is suitable for research and development in:

  • Automatic Speech Recognition (ASR)
  • Speaker diarization
  • Spoken dialogue systems
  • Dialogue analysis
  • Turn-taking analysis
  • Japanese speech corpora
  • ELAN-based annotation research
  • Speech segmentation and alignment
  • Tourism-domain conversational AI

Commercial Support

IngCrowd provides custom Japanese speech-data collection and annotation services, including:

  • Dialogue scenario design
  • Participant recruitment
  • Remote or on-site recording
  • Mixed and speaker-separated audio production
  • Manual verbatim transcription
  • Human-created ELAN annotation
  • Time-aligned speech segmentation
  • Audio, text, and video dataset production

For collaboration, custom dataset production, additional data, or corresponding video recordings, please contact us via email.


License

This dataset is released under the CC BY-NC 4.0 license.

For commercial licensing, additional data, or custom dataset production, please contact the dataset provider.


日本語

概要

Japanese Travel Agency Roleplay Dialogue Corpus は、旅行代理店のスタッフと顧客との相談場面を想定した、日本語によるロールプレイ対話10件を収録した対話コーパスです。

各対話には、2話者ミックス音声、話者別音声、および人手で作成した時間同期ELAN(.eaf)アノテーションを収録しています。

対話では、旅行先の選定、旅行日程の相談、交通手段、宿泊施設、観光地、食事など、旅行代理店で想定されるさまざまな相談内容を扱っています。本データは音声・対話研究およびデータ制作を目的として収録したロールプレイであり、実際の旅行代理店における接客や相談を録音したものではありません。

本リポジトリは、株式会社イングクラウドが提供する日本語対話データ収集、人手による書き起こし、ELANを用いた時間同期アノテーション、および音声データ制作のサンプルとして公開しています。


公開データ

本データセットには以下のデータを収録しています。

  • 日本語旅行代理店相談対話:10件
  • 2話者ミックス音声:10本
  • 話者別音声:20本
  • 人手で作成したELAN(.eaf)アノテーション:10本
  • 総収録時間:約126分

音声仕様

  • WAV形式
  • サンプリング周波数:16 kHz
  • 量子化ビット数:16 bit
  • モノラル

対応する映像データは本リポジトリには含まれていません。

研究目的で利用を希望される場合は、お問い合わせいただければ、提供可能な範囲で個別にご案内いたします。


特長

  • 旅行代理店スタッフと顧客による日本語ロールプレイ対話
  • 2話者ミックス音声および話者別音声
  • 人手で作成したELANアノテーション
  • 人手による発話忠実な書き起こし
  • 発話単位で時間同期されたアノテーション
  • フィラー、言いよどみ、繰り返し、笑い、聞き取り困難な発話など自然発話の特徴を保持
  • 行き先、交通手段、宿泊施設、観光計画など多様な旅行相談シナリオを収録
  • 音声認識、対話システム、日本語音声研究などに利用可能

データ収集

対話はZoomを用いたオンラインロールプレイとして収録しました。

各対話は、旅行代理店受付役1名と顧客役1名によるオンライン会話で構成されています。

参加者はZoomミーティングに参加し、あらかじめ設定された旅行代理店のシナリオに沿って自然な対話を行いました。

シナリオは以下のとおりです。

  • 顧客が旅行代理店を訪問する。
  • 旅行について相談する。
  • 受付担当者が旅行商品やプランを提案する。

収録シナリオ

ID 行き先・相談内容 スタッフ役 顧客役 収録時間
01 沖縄・那覇:観光施設、空港アクセス、レンタカー 40代女性 20代男性 14:07
02 静岡・熱海:両親・祖父母を含む夏の家族旅行 40代女性 20代男性 12:16
03 神奈川・箱根:温泉・宿泊・観光を含む家族旅行 40代女性 40代女性 12:16
04 千葉:家族3名での日帰り観光 40代女性 40代女性 18:41
05 茨城:祖父との春旅行 40代女性 20代男性 13:30
06 新潟:親との旅行に向けた観光相談 40代女性 20代男性 12:23
07 北海道・小樽:友人3名での旅行 40代女性 30代女性 11:31
08 沖縄・八重山諸島:母との夏旅行 40代女性 30代女性 11:27
09 東京:海外から来日する友人への観光案内 40代女性 30代女性 13:04
10 鹿児島:春の一人旅 40代女性 30代女性 12:19

参加者の情報は年代・性別のみを掲載しており、個人を特定できる情報は含まれていません。


データ構成

各対話フォルダには以下のファイルが含まれています。

japanese-roleplay-travel-agencyXX/
├── japanese-roleplay-travel-agencyXX_audio.wav
├── japanese-roleplay-travel-agencyXX_individual_audio1.wav
├── japanese-roleplay-travel-agencyXX_individual_audio2.wav
└── japanese-roleplay-travel-agencyXX.eaf
  • _audio.wav:2話者ミックス音声
  • _individual_audio1.wav、_individual_audio2.wav:話者別音声
  • .eaf:人手で作成したELAN形式の時間同期アノテーション

追加データ

本データセットには音声ファイルおよびELANアノテーションファイルを収録しています。 対応する映像データは、本リポジトリには含まれていません。 研究目的で対応する映像データの利用を希望される場合は、メール でお問い合わせください。


書き起こし・アノテーション

書き起こしはすべて人手によりELANを用いて作成しています。

各発話は人手で区切り、音声との時間対応を付与しています。

読みやすい文章へ整形するのではなく、フィラー、言いよどみ、繰り返し、笑い、聞き取りが困難な発話を含め、実際の発話内容をできる限り忠実に転記しています。

書き起こしおよびアノテーションには、自動音声認識(ASR)は使用していません。

主なTierは以下の2種類です。

  • staff:旅行代理店スタッフ
  • customer:旅行相談を行う顧客

想定用途

本データセットは以下の研究・開発用途を想定しています。

  • 日本語音声認識(ASR)
  • 話者分離
  • 音声対話システム
  • 対話分析
  • ターンテイキング分析
  • 日本語音声コーパス研究
  • ELANアノテーション研究
  • 音声区間分割・時間同期
  • 観光・旅行分野の対話AI

カスタムデータセット制作

株式会社イングクラウドでは、プロジェクトの要件に応じた日本語音声データセットの制作を行っています。

対応内容の例:

  • 対話シナリオ設計
  • 参加者募集・収録進行
  • 対面・リモート収録
  • 2話者ミックス音声・話者別音声作成
  • 人手による発話忠実な書き起こし
  • 人手によるELANアノテーション
  • 発話区間の時間同期
  • 音声・テキスト・映像データセット制作

カスタムデータセット制作、追加データ、対応する映像データについては、 メール でお問い合わせください。


ライセンス

本データセットは CC BY-NC 4.0 ライセンスで公開しています。


Custom Dataset Services / データセット制作のご相談

IngCrowd provides custom Japanese participant recruitment, speech and dialogue recording, transcription, time-aligned ELAN annotation, data collection, and human evaluation for AI development and academic research.

For custom dataset creation or research support, please visit our website or contact us by email.

株式会社イングクラウドでは、AI開発・学術研究向けに、研究参加者や話者の募集、日本語の音声・対話収録、文字起こし、ELANアノテーション、データ収集、人手評価などを提供しています。

独自のデータセット制作や研究支援については、ウェブサイトまたはメールよりお問い合わせください。

Downloads last month
163