Logo
Back to Blog
Fundamentals

Integrate the Pronunciation Assessment Android SDK in 10 Minutes

Run a DolphinSOE Android demo from trial credentials and AAR setup through engine initialization, recording or file assessment, and result callbacks.

Integrate the Pronunciation Assessment Android SDK in 10 Minutes

If you are planning to develop a language-learning app, once you understand what DolphinSOE can do, you will probably want to start building right away. DolphinSOE provides SDKs for multiple programming languages to help developers integrate Pronunciation Assessment quickly. Using the Android SDK as an example, this article walks you through getting an English Pronunciation Assessment demo that returns scores up and running in just 10 minutes.

DolphinSOE Pronunciation Assessment demo running in Android StudioDolphinSOE Android demo

Does integrating Pronunciation Assessment on Android really take only 10 minutes?

Running a minimum viable demo really does take only 10 minutes: apply for credentials, add the SDK, initialize the engine, record once, and receive a score. However, running the demo does not mean it is ready for production. A production launch still requires question-bank preparation, product UI development, exception handling, and other work.

Even for a demo, though, the process is much simpler than most people expect. The DolphinSOE Android SDK wraps the entire flow—including audio capture, token maintenance, audio transfer, and score retrieval—into just a few client-side calls. Starting from scratch, the process is roughly as follows:

StepEstimated time
Submit the website form and receive the automated reply2 minutes
Download the SDK and enter the appid/appsecret in the sample code2 minutes
Configure the assessment parameters, including the mode and reference answer2 minutes
Build and run4 minutes
Japanese diagram of the 10-minute integration flowIntegration flow in JapaneseChinese diagram of the 10-minute integration flowIntegration flow in Chinese

Prepare your development environment, and let’s get started.

Apply for trial access in two minutes

The DolphinSOE trial process is straightforward:

  • ① Submit a trial request: Complete the Contact Us form on the official website.
  • ② Receive the automated reply: Within a few seconds, the system automatically sends an email containing trial appid / appsecret credentials. These credentials support English, Japanese, and Chinese Pronunciation Assessment, with a Concurrent Connection limit of two—enough for local debugging and a single-device demonstration. The email also includes the SDK download URL and API documentation. Download and extract the Android SDK, and it is ready to use.

The downloaded demo project includes a complete test interface. Core features such as recording, audio playback, and score display are already implemented, so developers do not need to build the UI from scratch, making integration and debugging more efficient.

Run your first Pronunciation Assessment flow in three steps

Using a basic English assessment mode as an example, the demo reduces integration to three steps:

Step 1 · Add the SDK file. Place the .aar file from the archive in your project’s app/libs directory. Select the corresponding aar package for your target device architecture (x86 / ARM).

Step 2 · Enter the credentials. Open MainActivity.java in the project and enter the appid / appsecret from the email in these two constants:

// MainActivity.java
private static final String appIdEn = "";     // ← Enter the appid from the email
private static final String appSecretEn = ""; // ← Enter the appsecret from the email

Step 3 · Build and run. Run the demo, select an assessment mode on the test screen, start recording, read the prompt, and stop recording. The score appears immediately.

At this point, you have already run the simplest Pronunciation Assessment demo. It really is that straightforward.

If you want to embed the Pronunciation Assessment SDK in your own app, keep reading.

What does the demo do behind the scenes?

If you want to embed assessment in your own product UI instead of using the demo test page, you need to understand the engine initialization flow. You can find it in EvalLogic.java.

// 1) Create the engine singleton
SpeechEval eval = SpeechEval.createInstance(this);

// 2) Basic configuration
eval.setSampleRate(16000);                       // The Sampling Rate is fixed at 16000
eval.setListener(evalListener);                  // Implement EventListener to receive result/error events
eval.setInitLanguageEn(SpeechEval.InitSettingOnline.OralOnLine); // Use online assessment mode

// 3) Initialize the singleton
eval.setServerAPI("api.soe.dolphin-ai.jp");
eval.init(appIdEn, appSecretEn, "user01");

The initialization result is returned through the engineInitState callback. The success code is "00000". Common failure codes include 13010 (network permission missing), 11011 (signature error), and 11013 (appId does not exist). During development and debugging, use the code together with the status-code table to troubleshoot the issue.

Which parameters does an assessment require?

Parameters for an assessment fall into two categories: common parameters and assessment-mode parameters.

  • Common parameters: Control how to connect, which language to use, and which scoring rules apply. These settings apply to the entire session and are independent of what is being assessed.
  • Assessment-mode parameters: Control what the current item assesses and how it is assessed. These settings apply to each assessment task or request, with the assessment mode and reference answer at their core.

When you run the demo, UI.java already configures common parameters such as the language langType and Sampling Rate sampleRate. To configure a parameter, use the corresponding set method provided by the SpeechEval class. Note that if you do not want audio data retained in the cloud, call setAudioUrl(false). If you need to store audio in the cloud, call setAudioUrl(true). The audio download URL is then returned with the score, and cloud-stored data is automatically deleted after 30 days.

Pass assessment-mode parameters through setParamsJson:

JSONObject json = new JSONObject();
json.put("mode", "word");              // Mode: word / sentence / chapter, etc.
json.put("refText", "good morning");  // Reference text used by the engine for scoring
eval.setParamsJson(json);

Assessment modes differ significantly in how reference text is structured and how granular the scoring is. We have covered the duration limits, refText, and scoring dimensions of each mode systematically in our article on basic English Pronunciation Assessment modes. For the complete parameter list, using the English word mode as an example, see the API documentation.

Finally, we recommend sending a userId with every assessment. You can define it yourself; the demo uses "user01". The userId is returned unchanged with the assessment result, which helps with troubleshooting and data traceability.

How do you start recording and receive the score?

Once configuration is complete, there are two ways to start an assessment:

  • Recording: Create a recorder with createRecorder(), then call start(recorder, startListener). When the user finishes speaking, call stop() to end the recording. You can also enable trailing-silence detection to stop automatically. The startListener monitors the start and end states for a more stable experience.
  • File upload: Call start(wavPath, startListener) to assess an existing audio file directly. The file must be WAV or PCM at 16000 Hz / 16 Bit / mono. The task ends automatically after the file is sent, so you do not need to call stop().

To abort an assessment in progress, call cancel(). The canceled assessment will not return a result.

The result is returned as a JSON string in the onResult callback:

eval.setListener(new SpeechEval.EventListener() {
    @Override
    public void onResult(String result, boolean online) {
        // result is a JSON string; parse the scores for each dimension from it
        // See the API protocol documentation for the exact format
    }
    @Override
    public void onWarning(String taskId, String code, String msg) {
        // Handle Warning events
    }
    @Override
    public void onError(String taskId, String code, String msg) {
        // Handle Error events
    }
});

Callbacks are divided into Warning and Error. A Warning still returns an assessment result—for example, when the volume is too low or the recording is too short. An Error means the task has terminated abnormally and no score is returned—for example, when a required parameter is missing or the credentials are invalid.

Chinese flowchart showing Warning, Error, and assessment-result outcomesResult callback flow in ChineseJapanese flowchart showing Warning, Error, and assessment-result outcomesResult callback flow in Japanese

Conclusion

At this point, you should understand the basic DolphinSOE Android Pronunciation Assessment SDK integration flow and be able to run a complete Pronunciation Assessment demo quickly.

DolphinSOE Pronunciation Assessment has been deployed and validated in overseas teaching scenarios. Customers include language-learning apps, schools, after-school programs, and other educational organizations. The service supports approximately 100,000 assessment calls per day (100k calls/day), and has proven itself under real-world, high-concurrency teaching workloads.

SDK integration is only the first step. To realize the full value of Pronunciation Assessment, you also need to design practice flows around your own product and learners, turning scoring capabilities into a practical pronunciation-training experience. If you encounter problems during integration, consult the supporting API documentation or contact us for technical support.

Share Article