---
title: Automatic Speech Recognition
slug: automatic-speech-recognition
docTags: 
createdAt: 2024-04-05T07:39:06.269Z
---

The below input parameters are for different attack types. To start working with the APIs, see [Audio Speech Recognition](docId\:NlGwGbtjTaOqRQO8ITMya).

:::hint{type="warning"}
Audio Speech Recognition is an early access feature with limited functionality. It is not available as part of AIShield pypi package. For early access, kindly contact **AIShield.Contact\@bosch.com**
:::

-

:::hint{type="info"}
**ASR Models Support**:

1. [OpenAI-Whisper](https://github.com/openai/whisper/blob/main/model-card.md)**:&#x20;**&#x41;ll variants in PyTorch and Hugging Face formats are supported. Refer to the OpenAI-Whisper model card for the complete list of available variants. \[Ref: [OpenAI-Whisper-ModelCard](https://github.com/openai/whisper/blob/main/model-card.md)]
:::



# Common Parameter

The below table parameters are common for Evasion Attack type.

| Parameter                      | Data type | Description                                                                                                    | Remark                                                                                                                                                        |
| ------------------------------ | --------- | -------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| model\_Id                      | String    | Model\_id received during model registration.<br />We need to provide this model ID in query parameter in URL. | You have to do model registration only once for a model to perform model analysis. This will help you track the no of api call made, and it's success metric. |
| **Request Body (Json format)** |           |                                                                                                                |                                                                                                                                                               |
| model\_framework               | String    | Framework on which model is trained on.                                                                        | curretly supported framework are: **onnx&#x20;**&#x66;or **Evasion&#x20;**&#x61;nd **pytorch&#x20;**&#x66;or **data-poisoning** .                             |



:::hint{type="info"}
To access all sample artifacts, please visit [Artifacts](docId\:iJNEOCXoStabvvrsq11fa).&#x20;

- For specific artifact details,  refer&#x20;
  - Vulnerability Report : [Vulnerability Report](docId\:hL0uT2MWlCBkt8F97fr-W)    &#x20;
  - Sample Attacks : [Sample Attacks](docId:4G1mjM5lQjfm8t5WBVwpr)
:::





::::ExpandableHeading
# Evasion

## File upload format

- &#x20;**Data**: The processed audio data, ready to be passed to the model for prediction, should be saved in a folder.&#x20;

[Download sample data](https://aisdocs.blob.core.windows.net/reference/upload/Audio/AudioSpeechRecognition/Evasion/data.zip)

- **Label**: A CSV file should be created with two columns: "audio" and "label." The first column should contain the audio file name, and the second column should contain the label(text).  Check sample label file attached.

[Download sample label](https://aisdocs.blob.core.windows.net/reference/upload/Audio/AudioSpeechRecognition/Evasion/label.zip)

- **Model**: The model should be saved in onnx format. This can be ignored when model is hosted as an API.

[Download sample model](https://aisdocs.blob.core.windows.net/reference/upload/Audio/AudioSpeechRecognition/Evasion/model.zip)

:::hint{type="warning"}
**Note**:

1. **Format:** All uploaded files must be in a zipped format.
2. **Dataset:** The provided files are sample audio data from the Librespeech dataset.
3. **Audio File Properties:** Each audio file should not exceed 30 seconds and must have a sampling rate of 16,000 Hz.
4. **Sample Count:** The minimum number of samples required must be between 300 and 500.
5. **Data Sampling Strategy:** Ensure that the uploaded samples are representative of the complete dataset, using techniques such as normal distributive sampling for even distribution.
:::

## &#x20;Parameter

| Parameter                      | Data type | Description                                                                            | Remark                                                                                                                                          |
| ------------------------------ | --------- | -------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------- |
| **Request Body (Json format)** |           |                                                                                        |                                                                                                                                                 |
| model\_api\_details            | String    | Provide API details of hosted model as encrypted JSON string <br />                    | provide this only if use\_model\_api is "yes".                                                                                                  |
| use\_model\_api                | String    | Use model API to use your model endpoint instead of uploading the model as a zip file. | when this parameter is yes, you don't have to upload model as zip. You can pass api url along with other verification credential in json file.  |

## Convert pytorch model to ONNX format for ASR

Converting Pytorch model to onnx will create two models (Encoder and Decoder) in onnx format.&#x20;

- To Load an OpenAI Model:
  - &#x20; load\_model: "tiny.en", "base.en", "small.en", "medium.en", "large.en"&#x20;
- To Load a Distil-Whisper Model:
  - load\_model: model.bin or "distil-whisper/distil-large-v3"

### Loading of Whisper model

```python
import whisper, torch, onnxruntime

# load tiny model trained on whisper dataset
device = torch.device('cpu')
model = whisper.load_model('tiny.en').to(device) #Hugging Face --> Model.bin
tokenizer = whisper.decoding.get_tokenizer(False,language='en',task='transcribe') ##initialoze whisper tokenizer
model.requires_grad_(False)
model.eval()
```

### Convert to Onnx Format

**&#x20;First load sample audio file to get the feature of audio signal:**

:::CodeblockTabs
python

```python
import librosa

audio = librosa.load(path_of_audio_file,sr=16000)[0] #sample audio file to load audio features
audio = whisper.pad_or_trim(audio)
log_mel = whisper.log_mel_spectrogram(audio).unsqueeze(axis=0)
```
:::

**Prepare the input for decoder model:&#x20;**

:::CodeblockTabs
python

```python
input = log_mel.to(device)
audio_features = model.encoder(input)
dec_tokens = torch.tensor([tokenizer.sot_sequence_including_notimestamps],dtype=torch.long) ##decoder token

```
:::

**Convert and save model (encoder and decoder) in onnx format:**

:::CodeblockTabs
python

```python

#: ONNX export
##Encoder Export
torch.onnx.export(
    model.encoder, 
    (log_mel,), 
    "distil_encoder.onnx", 
    input_names=["input_mel"], 
    output_names=["out"],
    dynamic_axes={
        "input_mel": {0: "batch"},
        "out": {0: "batch"},
    },
)

#Decoder Export
torch.onnx.export(
    model.decoder, 
    (dec_tokens, audio_features), 
    "distil_decoder.onnx", 
    input_names=["tokens", "audio"], 
    output_names=["out"], 
    dynamic_axes={
        "tokens": {0: "batch", 1: "seq"},
        "audio": {0: "batch"},
        "out": {0: "batch", 1: "seq"},
    },
)
```
:::
::::

::::ExpandableHeading
# Data-Poisoning

## File upload format

- &#x20;**Reference Data**: The processed clean audio data, ready to be passed to the model for prediction, should be saved in a folder.&#x20;

[Download sample data](https://aisdocs.blob.core.windows.net/reference/upload/Audio/AudioSpeechRecognition/Data-Poisoning/data.zip)

- **Refernce Label**: A CSV file corresponding to clean data should be created with two columns: "audio" and "label." The first column should contain the audio file name, and the second column should contain the label(text).  Check sample label file attached.

[Download sample label](https://aisdocs.blob.core.windows.net/reference/upload/Audio/AudioSpeechRecognition/Data-Poisoning/label.zip)

- **Model**: The model should be saved in onnx format. This can be ignored when model is hosted as an API.

[Download sample model](https://aisdocs.blob.core.windows.net/reference/upload/Audio/AudioSpeechRecognition/Data-Poisoning/model.zip)

- &#x20;**Universal Data**: Audio samples under test (poisoned, clean samples), should be saved in a folder .&#x20;

[Download sample universal data](https://aisdocs.blob.core.windows.net/reference/upload/Audio/AudioSpeechRecognition/Data-Poisoning/universal_data.zip)

- &#x20;**Universal Label**:  A CSV files corresponding to  universal data hould be created with two columns: "audio" and "label." The first column should contain the audio file name, and the second column should contain the label(text).  Check sample label file attached.

[Download sample universal label](https://aisdocs.blob.core.windows.net/reference/upload/Audio/AudioSpeechRecognition/Data-Poisoning/universal_label.zip)



:::hint{type="warning"}
**Note**:

1. **Format:** All uploaded files must be in a zipped format.
2. **Dataset:** The provided files are sample audio data from the Librespeech dataset.
3. **Audio File Properties:** Each audio file should not exceed 30 seconds and must have a sampling rate of 16,000 Hz.
4. **Sample Count:** The minimum number of samples required must be between 500 to 700.
5. **Data Sampling Strategy:** Ensure that the uploaded samples are representative of the complete dataset, using techniques such as normal distributive sampling for even distribution.
:::

## Conversion of OpenAI-Whisper models to HuggingFace format



::githubGist{url="https://gist.github.com/c4f0feb9fbb5a79d0726fc8baa2cc305.js"}
::::

&#x20;
