New needs for VoiceAI preprocessing by old tools

Stuart Naylor <[email protected]> Thu, 20 Aug 2020 07:07:47 +0000
Newsgroups gmane.comp.audio.compression.speex.devel
Message-ID <DB7P191MB03322B78E8A8A59D3679EE8FA85A0@DB7P191MB0332.EURP191.PROD.OUTLOOK.COM>
--===============3788620128487000883==
Content-Language: en-GB
Content-Type: multipart/alternative;
	boundary="_000_DB7P191MB03322B78E8A8A59D3679EE8FA85A0DB7P191MB0332EURP_"

--_000_DB7P191MB03322B78E8A8A59D3679EE8FA85A0DB7P191MB0332EURP_
Content-Type: text/plain; charset="Windows-1252"
Content-Transfer-Encoding: quoted-printable

Hi,

Xiph has done some really good work but since Speex things have seemed quit=
e quiet.
Devs to hobbyists have a huge battle with VoiceAI audio processing and even=
 when available extremely similar routines are separated across libs so tha=
t much process duplication occurs and increases load.

The Alsa-plugins of Speex echo just don=92t seem to work whilst repos using=
 the speex libs do.
https://github.com/voice-engine/ec

I actually use the speex AGC Alsa-plugin as its a great tool and then run E=
cho via the above.
Coupled to WebRTC_VAD and a python based MFCC routine.
Beamforming and DOA I have given up hope for as I am running so many extrem=
ely similar FFT routines concurrently in different process space that my lo=
ad usage makes me believe there isn=92t a chance.

All of them have extremely similar frame based spectra analysis collection =
that collate and process over different windows for different process but m=
uch of the input process are almost identical heavy load FFT process that f=
or some reason we duplicate because the Linux landscape is one of scattered=
 audio processing libs.

Linux currently has much work in terms of open source ASR, TTS, TensorFlow =
frameworks but what we have available in terms of audio pre-processing are =
scattered individual segments of the audio chain and create a seriously ine=
fficient chain.
These are old tools but the whole chain to MFCC output just doesn=92t exist=
.

Stuart

Sent from Mail<https://go.microsoft.com/fwlink/?LinkId=3D550986> for Window=
s 10


--_000_DB7P191MB03322B78E8A8A59D3679EE8FA85A0DB7P191MB0332EURP_
Content-Type: text/html; charset="Windows-1252"
Content-Transfer-Encoding: quoted-printable

<html xmlns:o=3D"urn:schemas-microsoft-com:office:office" xmlns:w=3D"urn:sc=
hemas-microsoft-com:office:word" xmlns:m=3D"http://schemas.microsoft.com/of=
fice/2004/12/omml" xmlns=3D"http://www.w3.org/TR/REC-html40">
<head>
<meta http-equiv=3D"Content-Type" content=3D"text/html; charset=3DWindows-1=
252">
<meta name=3D"Generator" content=3D"Microsoft Word 15 (filtered medium)">
<style><!--
/* Font Definitions */
@font-face
	{font-family:"Cambria Math";
	panose-1:2 4 5 3 5 4 6 3 2 4;}
@font-face
	{font-family:Calibri;
	panose-1:2 15 5 2 2 2 4 3 2 4;}
/* Style Definitions */
p.MsoNormal, li.MsoNormal, div.MsoNormal
	{margin:0cm;
	font-size:11.0pt;
	font-family:"Calibri",sans-serif;}
a:link, span.MsoHyperlink
	{mso-style-priority:99;
	color:blue;
	text-decoration:underline;}
.MsoChpDefault
	{mso-style-type:export-only;}
@page WordSection1
	{size:612.0pt 792.0pt;
	margin:72.0pt 72.0pt 72.0pt 72.0pt;}
div.WordSection1
	{page:WordSection1;}
--></style>
</head>
<body lang=3D"EN-GB" link=3D"blue" vlink=3D"#954F72">
<div class=3D"WordSection1">
<p class=3D"MsoNormal">Hi,</p>
<p class=3D"MsoNormal"><br>
Xiph has done some really good work but since Speex things have seemed quit=
e quiet.<br>
Devs to hobbyists have a huge battle with VoiceAI audio processing and even=
 when available extremely similar routines are separated across libs so tha=
t much process duplication occurs and increases load.<br>
<br>
The Alsa-plugins of Speex echo just don=92t seem to work whilst repos using=
 the speex libs do.<br>
<a href=3D"https://github.com/voice-engine/ec">https://github.com/voice-eng=
ine/ec</a><br>
<br>
I actually use the speex AGC Alsa-plugin as its a great tool and then run E=
cho via the above.<br>
Coupled to WebRTC_VAD and a python based MFCC routine.<br>
Beamforming and DOA I have given up hope for as I am running so many extrem=
ely similar FFT routines concurrently in different process space that my lo=
ad usage makes me believe there isn=92t a chance.<br>
<br>
All of them have extremely similar frame based spectra analysis collection =
that collate and process over different windows for different process but m=
uch of the input process are almost identical heavy load FFT process that f=
or some reason we duplicate because
 the Linux landscape is one of scattered audio processing libs.<br>
<br>
Linux currently has much work in terms of open source ASR, TTS, TensorFlow =
frameworks but what we have available in terms of audio pre-processing are =
scattered individual segments of the audio chain and create a seriously ine=
fficient chain.<br>
These are old tools but the whole chain to MFCC output just doesn=92t exist=
.<br>
<br>
Stuart<br>
<br>
</p>
<p class=3D"MsoNormal">Sent from <a href=3D"https://go.microsoft.com/fwlink=
/?LinkId=3D550986">
Mail</a> for Windows 10</p>
<p class=3D"MsoNormal"><o:p>&nbsp;</o:p></p>
</div>
</body>
</html>

--_000_DB7P191MB03322B78E8A8A59D3679EE8FA85A0DB7P191MB0332EURP_--

--===============3788620128487000883==
Content-Type: text/plain; charset="utf-8"
MIME-Version: 1.0
Content-Transfer-Encoding: base64
Content-Disposition: inline

X19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX18KU3BlZXgtZGV2
IG1haWxpbmcgbGlzdApTcGVleC1kZXZAeGlwaC5vcmcKaHR0cDovL2xpc3RzLnhpcGgub3JnL21h
aWxtYW4vbGlzdGluZm8vc3BlZXgtZGV2Cg==

--===============3788620128487000883==--