Integrate the Pronunciation Assessment Android SDK in 10 Minutes
Run a DolphinSOE Android demo from trial credentials and AAR setup through engine initialization, recording or file assessment, and result callbacks.

If you are planning to develop a language-learning app, once you understand what DolphinSOE can do, you will probably want to start building right away. DolphinSOE provides SDKs for multiple programming languages to help developers integrate Pronunciation Assessment quickly. Using the Android SDK as an example, this article walks you through getting an English Pronunciation Assessment demo that returns scores up and running in just 10 minutes.
DolphinSOE Android demoDoes integrating Pronunciation Assessment on Android really take only 10 minutes?
Running a minimum viable demo really does take only 10 minutes: apply for credentials, add the SDK, initialize the engine, record once, and receive a score. However, running the demo does not mean it is ready for production. A production launch still requires question-bank preparation, product UI development, exception handling, and other work.
Even for a demo, though, the process is much simpler than most people expect. The DolphinSOE Android SDK wraps the entire flow—including audio capture, token maintenance, audio transfer, and score retrieval—into just a few client-side calls. Starting from scratch, the process is roughly as follows:
| Step | Estimated time |
|---|---|
| Submit the website form and receive the automated reply | 2 minutes |
| Download the SDK and enter the appid/appsecret in the sample code | 2 minutes |
| Configure the assessment parameters, including the mode and reference answer | 2 minutes |
| Build and run | 4 minutes |
Integration flow in Japanese
Integration flow in ChinesePrepare your development environment, and let’s get started.
Apply for trial access in two minutes
The DolphinSOE trial process is straightforward:
- ① Submit a trial request: Complete the Contact Us form on the official website.
- ② Receive the automated reply: Within a few seconds, the system automatically sends an email containing trial
appid/appsecretcredentials. These credentials support English, Japanese, and Chinese Pronunciation Assessment, with a Concurrent Connection limit of two—enough for local debugging and a single-device demonstration. The email also includes the SDK download URL and API documentation. Download and extract the Android SDK, and it is ready to use.
The downloaded demo project includes a complete test interface. Core features such as recording, audio playback, and score display are already implemented, so developers do not need to build the UI from scratch, making integration and debugging more efficient.
Run your first Pronunciation Assessment flow in three steps
Using a basic English assessment mode as an example, the demo reduces integration to three steps:
Step 1 · Add the SDK file. Place the .aar file from the archive in your project’s app/libs directory. Select the corresponding aar package for your target device architecture (x86 / ARM).
Step 2 · Enter the credentials. Open MainActivity.java in the project and enter the appid / appsecret from the email in these two constants:
// MainActivity.java private static final String appIdEn = ""; // ← Enter the appid from the email private static final String appSecretEn = ""; // ← Enter the appsecret from the email
Step 3 · Build and run. Run the demo, select an assessment mode on the test screen, start recording, read the prompt, and stop recording. The score appears immediately.
At this point, you have already run the simplest Pronunciation Assessment demo. It really is that straightforward.
If you want to embed the Pronunciation Assessment SDK in your own app, keep reading.
What does the demo do behind the scenes?
If you want to embed assessment in your own product UI instead of using the demo test page, you need to understand the engine initialization flow. You can find it in EvalLogic.java.
// 1) Create the engine singleton SpeechEval eval = SpeechEval.createInstance(this); // 2) Basic configuration eval.setSampleRate(16000); // The Sampling Rate is fixed at 16000 eval.setListener(evalListener); // Implement EventListener to receive result/error events eval.setInitLanguageEn(SpeechEval.InitSettingOnline.OralOnLine); // Use online assessment mode // 3) Initialize the singleton eval.setServerAPI("api.soe.dolphin-ai.jp"); eval.init(appIdEn, appSecretEn, "user01");
The initialization result is returned through the engineInitState callback. The success code is "00000". Common failure codes include 13010 (network permission missing), 11011 (signature error), and 11013 (appId does not exist). During development and debugging, use the code together with the status-code table to troubleshoot the issue.
Which parameters does an assessment require?
Parameters for an assessment fall into two categories: common parameters and assessment-mode parameters.
- Common parameters: Control how to connect, which language to use, and which scoring rules apply. These settings apply to the entire session and are independent of what is being assessed.
- Assessment-mode parameters: Control what the current item assesses and how it is assessed. These settings apply to each assessment task or request, with the assessment mode and reference answer at their core.
When you run the demo, UI.java already configures common parameters such as the language langType and Sampling Rate sampleRate. To configure a parameter, use the corresponding set method provided by the SpeechEval class. Note that if you do not want audio data retained in the cloud, call setAudioUrl(false). If you need to store audio in the cloud, call setAudioUrl(true). The audio download URL is then returned with the score, and cloud-stored data is automatically deleted after 30 days.
Pass assessment-mode parameters through setParamsJson:
JSONObject json = new JSONObject(); json.put("mode", "word"); // Mode: word / sentence / chapter, etc. json.put("refText", "good morning"); // Reference text used by the engine for scoring eval.setParamsJson(json);
Assessment modes differ significantly in how reference text is structured and how granular the scoring is. We have covered the duration limits, refText, and scoring dimensions of each mode systematically in our article on basic English Pronunciation Assessment modes. For the complete parameter list, using the English word mode as an example, see the API documentation.
Finally, we recommend sending a userId with every assessment. You can define it yourself; the demo uses "user01". The userId is returned unchanged with the assessment result, which helps with troubleshooting and data traceability.
How do you start recording and receive the score?
Once configuration is complete, there are two ways to start an assessment:
- Recording: Create a recorder with
createRecorder(), then callstart(recorder, startListener). When the user finishes speaking, callstop()to end the recording. You can also enable trailing-silence detection to stop automatically. ThestartListenermonitors the start and end states for a more stable experience. - File upload: Call
start(wavPath, startListener)to assess an existing audio file directly. The file must be WAV or PCM at 16000 Hz / 16 Bit / mono. The task ends automatically after the file is sent, so you do not need to callstop().
To abort an assessment in progress, call cancel(). The canceled assessment will not return a result.
The result is returned as a JSON string in the onResult callback:
eval.setListener(new SpeechEval.EventListener() { @Override public void onResult(String result, boolean online) { // result is a JSON string; parse the scores for each dimension from it // See the API protocol documentation for the exact format } @Override public void onWarning(String taskId, String code, String msg) { // Handle Warning events } @Override public void onError(String taskId, String code, String msg) { // Handle Error events } });
Callbacks are divided into Warning and Error. A Warning still returns an assessment result—for example, when the volume is too low or the recording is too short. An Error means the task has terminated abnormally and no score is returned—for example, when a required parameter is missing or the credentials are invalid.
Result callback flow in Chinese
Result callback flow in JapaneseConclusion
At this point, you should understand the basic DolphinSOE Android Pronunciation Assessment SDK integration flow and be able to run a complete Pronunciation Assessment demo quickly.
DolphinSOE Pronunciation Assessment has been deployed and validated in overseas teaching scenarios. Customers include language-learning apps, schools, after-school programs, and other educational organizations. The service supports approximately 100,000 assessment calls per day (100k calls/day), and has proven itself under real-world, high-concurrency teaching workloads.
SDK integration is only the first step. To realize the full value of Pronunciation Assessment, you also need to design practice flows around your own product and learners, turning scoring capabilities into a practical pronunciation-training experience. If you encounter problems during integration, consult the supporting API documentation or contact us for technical support.
Share Article
Read more

Why Children Need a Dedicated Kids Model for English Pronunciation Assessment
Explore why children need a dedicated English Pronunciation Assessment model, how DolphinSOE Kids models calibrate scores, and how to switch task modes.

Nearly All 26,000 Patient Visit Records Contained Fabricated Content: The Structural Risks of Whisper Hallucinations and Business Countermeasures
Why does Whisper fabricate entire sentences during silence? Drawing on the AP investigation and Cornell University research, this article explains hallucinations in LLM decoders, their cost in four business settings, why post-processing cannot solve them, and a practical silence-testing checklist.

From Phonemes to Paragraphs: Understanding the Basic English Pronunciation Assessment Question Types
How should you choose among DolphinSOE phoneme, word, sentence, chapter, and correction question types? Compare their granularity, limits, scoring dimensions, and best use cases.