Choosing Between the Three Products
Differences between AI Translation, AI Speaking and AI Speaking, plus common questions.
Choose in one sentence
| Product | Who it suits | In one sentence |
|---|---|---|
| AI Translation | Simultaneous interpreting, following a meeting, listening to someone speak a foreign language | Listens to the other person, transcribes in real time and translates into a language you can read |
| AI Speaking | Practising a foreign language and your delivery | Listens to you, transcribes sentence by sentence and gives practice feedback |
| AI Speaking | Mock speakings, preparing answers | Records your answers, and in session 2 turns speaking questions into AI reference answers |
Need to understand / translate what "someone else" says in a foreign language?
�?AI Translation
Need to "speak it yourself" and get sentence-by-sentence feedback?
�?AI Speaking
Need a mock speaking: record your own answers + get reference scripts for questions?
�?AI Speaking
Comparing the two session areas
| AI Translation | AI Speaking | AI Speaking | |
|---|---|---|---|
| Session 1 listens to | the other person | you | you (the candidate) |
| Session 1 main output | Source text + translation | Transcript + practice feedback | Transcript + speaking feedback |
| Session 2 main use | Text questions / multilingual output | Text practice dialogue | Speaking question �?reference answer |
| Session 2 read-aloud | �? | �? | �?(by default) |
AI Translation
| Area | What you do | What the system does |
|---|---|---|
| Session 1 | Listen to the other person speaking a foreign language | Transcribes the source and translates it into your reading language |
| Session 2 | Ask questions / add context in text (optional) | Generates a reply/translation aimed at the other person; can be read aloud |
AI Speaking
| Area | What you do | What the system does |
|---|---|---|
| Session 1 | You speak into the microphone | Transcribes + gives practice feedback (not interpretation) |
| Session 2 | Practise in text (optional) | AI replies in text; can be read aloud |
The read-aloud button lives in session 2. Session 1 is transcription and feedback and does not play audio automatically.
AI Speaking
| Area | What you do | What the system does |
|---|---|---|
| Session 1 | Speak as the candidate in a simulation | Transcribes + gives speaking-oriented feedback |
| Session 2 | Enter speaking questions | Returns key points for a reference answer; no read-aloud by default |
Common misconceptions
| Misconception | What actually happens |
|---|---|
| Speaking = translation plus one extra chat | Session 1 in Speaking is already practice feedback, not interpretation |
| Speaking = translation with different wording | Session 1 in Speaking works like Speaking (you speak �?transcript + feedback); the business settings and session 2 make the difference |
| Session 1 replies with voice automatically in all three products | Read-aloud is mainly in session 2 (available in Translation and Speaking; off by default in Speaking) |
| Session 1 is "interpretation" in all three products | Only Translation listens to the other person and translates |
What to fill in under business settings
| Entry | Content |
|---|---|
| Settings �?YOYO assistant �?AI Language Partner | Auto-save interval, debug panel (client-side preferences) |
| Each product page �?business settings | Language direction, scene and vocabulary list, speaking toggles, etc. (feeds the AI pipeline) |
Examples of scene and vocabulary list (optional):
- Translation:
Broad context: simultaneous interpreting in a business meeting/Vocabulary: agenda, quote, contract - Speaking:
Broad context: everyday Japanese conversation practice, N3/Topics: travel, ordering food - Speaking:
Broad context: Japanese-language speaking for a backend engineer/Role/company: …
Filling in the scene and keywords improves transcription and terminology accuracy; you do not need to pick a "recording environment" separately in settings.
Time interval (seconds): real-time interpretation is segmented automatically by the system; this interval only applies when the fallback recording mode downgrades to chunking.
Common questions
Q: Why are there two session areas?
A: Session 1 handles the main voice flow (listening or speaking); session 2 handles text support (questions, practice dialogue, speaking questions).
Q: Why does only Translation show "source / translation"?
A: Only Translation listens to the other person, so it uses an interpretation-style display. Speaking and Speaking show what you said + AI feedback.
Q: Where is read-aloud?
A: In Translation and Speaking it sits next to the reply in session 2; Speaking does not offer read-aloud by default.
Q: Do I need to fill in the content summary?
A: It is optional. Filling in scene + keywords improves transcription and terminology accuracy and takes effect from the start of the session.
Q: How is this different from YOYO?
A: YOYO is a general assistant; the AI Language Partner is designed for long voice sessions, with chunking, an interpretation/feedback pipeline, bookings and metering.
How to use it (general)
- Open the relevant product page
- Complete the business settings
- Start the conversation �?record / interpret in session 1
- Optional: continue in text in session 2
- When you finish, you can save it as a voice session