July 2026 Newsletter - LDC

Penn LDC via Corpora <[email protected]> Wed, 15 Jul 2026 14:31:55 +0000
Newsgroups gmane.science.linguistics.corpora
Message-ID <BN0PR10MB50295919A09C53CF0FC6BCB3F9F82@BN0PR10MB5029.namprd10.prod.outlook.com>
--===============5788192565136502589==
Content-Language: en-US
Content-Type: multipart/alternative;
	boundary="_000_BN0PR10MB50295919A09C53CF0FC6BCB3F9F82BN0PR10MB5029namp_"

--_000_BN0PR10MB50295919A09C53CF0FC6BCB3F9F82BN0PR10MB5029namp_
Content-Type: text/plain; charset="us-ascii"
Content-Transfer-Encoding: quoted-printable

In this newsletter:
Fall 2026 LDC data scholarship program

New publications:
2012 NIST Speaker Recognition Evaluation Test Set<https://catalog.ldc.upenn=
.edu/LDC2026S09>
CALLHOME American English Second Edition<https://catalog.ldc.upenn.edu/LDC2=
026S08>
CALLHOME American English Lexicon (PRONLEX) Second Edition<https://catalog.=
ldc.upenn.edu/LDC2026L05>

________________________________
Fall 2026 LDC data scholarship program
Student applications for the Fall 2026 LDC data scholarship program are bei=
ng accepted now through September 15, 2026. This program provides eligible =
students with no-cost access to LDC data. Students must complete an applica=
tion consisting of a data use proposal and letter of support from their adv=
isor. For application requirements and program rules, visit the LDC Data Sc=
holarships<https://www.ldc.upenn.edu/language-resources/data/data-scholarsh=
ips> page.
________________________________

New publications:
2012 NIST Speaker Recognition Evaluation Test Set<https://catalog.ldc.upenn=
.edu/LDC2026S09> was developed by LDC and NIST and contains 10,321 hours of=
 English conversational telephone speech and in-person recorded studio sess=
ions for evaluation and modeling, along with answer keys, trial files, and =
documentation from the NIST-sponsored 2012 Speaker Recognition Evaluation (=
SRE12)<https://www.nist.gov/itl/iad/mltg/speaker-recognition-evaluation-201=
2>. SRE12 introduced a revised evaluation structure in which training data =
for target speakers was drawn from prior SRE corpora developed by LDC and w=
as provided in advance of the evaluation period.

Test data was drawn from Mixer 7 English Speech (LDC2025S08)<https://catalo=
g.ldc.upenn.edu/LDC2025S08> and REMIX Telephone Collection (LDC2023S09)<htt=
ps://catalog.ldc.upenn.edu/LDC2023S09>. Those datasets also provided segmen=
ts for modeling data; other modeling segments were drawn from Mixer 3 Speec=
h (LDC2023S02)<https://catalog.ldc.upenn.edu/LDC2023S02>, Mixer 4 and 5 Spe=
ech (LDC2020S03)<https://catalog.ldc.upenn.edu/LDC2020S03>, and Mixer 6 Spe=
ech (LDC2013S03)<https://catalog.ldc.upenn.edu/LDC2013S03>. The test data c=
ontains English speech only; some non-English speech is contained in modeli=
ng segments.

This release is comprised of 130,844 test segments, specifically, 83,778 ca=
ll segments and 47,066 interview segments. Modeling data consists of 46,948=
 segments.

2026 members can access this corpus through their LDC accounts. Non-members=
 may license this data for a fee.

*

CALLHOME American English Second Edition<https://catalog.ldc.upenn.edu/LDC2=
026S08> was developed by LDC and contains 56 hours of speech from 120 unscr=
ipted telephone conversations between native American English speakers. Thi=
s publication is a re-release of the original CALLHOME American English col=
lection, combining CALLHOME American English Speech (LDC97S42)<https://cata=
log.ldc.upenn.edu/LDC97S42> and CALLHOME American English Transcripts (LDC9=
7T14)<https://catalog.ldc.upenn.edu/LDC97T14>, with additional transcriptio=
n and updated directory structure, file formats, and documentation.

This release contains the 120 telephone conversations published in CALLHOME=
 American English Speech which represented training and development data an=
d a subset of evaluation data. Participants spoke on topics of their choice=
 in a single telephone call lasting up to 30 minutes. Calls were manually a=
udited for gender, language, recording quality, channel characteristics, di=
alect, and region. For this second edition, all audio was converted from SP=
HERE files to FLAC format, and the original training/development/evaluation=
 partitioning was removed.

This release also features revised transcripts conforming to updated LDC tr=
anscription guidelines that addressed normalization of annotation formats, =
standardization of speaker-produced and background noises, application of f=
oreign-language marking, whitespace cleanup, and corrections and consistenc=
y fixes.

The CALLHOME series consists of telephone conversations and transcripts dev=
eloped by LDC and Rutgers, The State University of New Jersey, in support o=
f research in speaker identification, language identification, and related =
technologies. Languages in the series include American English, Egyptian Ar=
abic, German, Japanese, Mandarin Chinese, and Spanish.

2026 members can access this corpus through their LDC accounts. Non-members=
 may license this data for a fee.

*

CALLHOME American English Lexicon (PRONLEX) Second Edition<https://catalog.=
ldc.upenn.edu/LDC2026L05> was developed by LDC and contains 90,988 English =
words with citation-form pronunciations. This second edition updates file f=
ormats, directory structure, and documentation. The first edition is availa=
ble as CALLHOME American English Lexicon (PRONLEX) (LDC97L20)<https://catal=
og.ldc.upenn.edu/LDC97L20>.

The words in the lexicon were derived from Wall Street Journal text used in=
 the continuous speech recognition publication series CSR-1 WSJ0 Complete (=
LDC93S6A<https://catalog.ldc.upenn.edu/LDC93S6A>), transcripts<https://isip=
.piconepress.com/projects/switchboard/> from the Switchboard telephone coll=
ection (LDC97S62)<https://catalog.ldc.upenn.edu/LDC97S62>, and transcripts =
representing unscripted telephone conversations between native American Eng=
lish speakers contained in CALLHOME American English Second Edition (LDC202=
6S08)<https://catalog.ldc.upenn.edu/LDC2026S08>.

PRONLEX transcription is a phonemic transcription system designed to suppor=
t speech recognition by providing a consistent and simplified representatio=
n of how words are pronounced in standard American English that allows vari=
ation to be generated later to avoid listing many pronunciation variations =
for each word. This single systematic base form can be expanded through rul=
es or modeling. The transcription was created using a modified ARPABET phon=
eme set<https://learnius.com/slp/3+Speech+Production%2C+Perception+and+Phon=
etics/4+Phonetics/2+Phonetic+Alphabet/ARPAbet>.

The lexicon contains three tab-separated information fields: (1) word: orth=
ographic representation of word; (2) pron: transcribed citation-form pronun=
ciations using modified ARPABET phoneme set; and (3) comments: (OPTIONAL) c=
omment on the entry. It is presented as a tab-delimited TSV file encoded in=
 UTF-8 format and includes a pronunciation dictionary derived from the lexi=
con in UTF-8 encoded CMUdict<https://stdlib.io/docs/api/latest/@stdlib/data=
sets/cmudict> format.

2026 members can access this corpus through their LDC accounts provided the=
y have submitted a completed copy of the special license agreement. Non-mem=
bers may license this data for a fee.

To unsubscribe from this newsletter, log in to your LDC account<https://cat=
alog.ldc.upenn.edu/login> and uncheck the box next to "Receive Newsletter" =
under Account Options or contact LDC for assistance.

Membership Coordinator
Linguistic Data Consortium<ldc.upenn.edu>
University of Pennsylvania
T: +1-215-573-1275
E: [email protected]<mailto:[email protected]>
M: 3600 Market St. Suite 810
      Philadelphia, PA 19104




--_000_BN0PR10MB50295919A09C53CF0FC6BCB3F9F82BN0PR10MB5029namp_
Content-Type: text/html; charset="us-ascii"
Content-Transfer-Encoding: quoted-printable

<html xmlns:v=3D"urn:schemas-microsoft-com:vml" xmlns:o=3D"urn:schemas-micr=
osoft-com:office:office" xmlns:w=3D"urn:schemas-microsoft-com:office:word" =
xmlns:m=3D"http://schemas.microsoft.com/office/2004/12/omml" xmlns=3D"http:=
//www.w3.org/TR/REC-html40">
<head>
<meta http-equiv=3D"Content-Type" content=3D"text/html; charset=3Dus-ascii"=
>
<meta name=3D"Generator" content=3D"Microsoft Word 15 (filtered medium)">
<!--[if !mso]><style>v\:* {behavior:url(#default#VML);}
o\:* {behavior:url(#default#VML);}
w\:* {behavior:url(#default#VML);}
.shape {behavior:url(#default#VML);}
</style><![endif]--><style><!--
/* Font Definitions */
@font-face
	{font-family:"Cambria Math";
	panose-1:2 4 5 3 5 4 6 3 2 4;}
@font-face
	{font-family:Aptos;}
/* Style Definitions */
p.MsoNormal, li.MsoNormal, div.MsoNormal
	{margin:0in;
	font-size:12.0pt;
	font-family:"Aptos",sans-serif;
	mso-ligatures:standardcontextual;}
a:link, span.MsoHyperlink
	{mso-style-priority:99;
	color:#467886;
	text-decoration:underline;}
span.EmailStyle17
	{mso-style-type:personal-compose;
	font-family:"Aptos",sans-serif;
	color:windowtext;}
.MsoChpDefault
	{mso-style-type:export-only;}
@page WordSection1
	{size:8.5in 11.0in;
	margin:1.0in 1.0in 1.0in 1.0in;}
div.WordSection1
	{page:WordSection1;}
--></style><!--[if gte mso 9]><xml>
<o:shapedefaults v:ext=3D"edit" spidmax=3D"1026" />
</xml><![endif]--><!--[if gte mso 9]><xml>
<o:shapelayout v:ext=3D"edit">
<o:idmap v:ext=3D"edit" data=3D"1" />
</o:shapelayout></xml><![endif]-->
</head>
<body lang=3D"EN-US" link=3D"#467886" vlink=3D"#96607D" style=3D"word-wrap:=
break-word">
<div class=3D"WordSection1">
<p class=3D"MsoNormal"><i><span style=3D"font-size:11.0pt">In this newslett=
er: <br>
</span></i><b><span style=3D"font-size:11.0pt">Fall 2026 LDC data scholarsh=
ip program
<o:p></o:p></span></b></p>
<p class=3D"MsoNormal"><b><span style=3D"font-size:11.0pt">&nbsp;&nbsp;&nbs=
p;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; =
</span></b><i><span style=3D"font-size:11.0pt"><br>
New publications:<br>
</span></i><span style=3D"font-size:11.0pt"><a href=3D"https://catalog.ldc.=
upenn.edu/LDC2026S09">2012 NIST Speaker Recognition Evaluation Test Set</a>=
<i><br>
</i><a href=3D"https://catalog.ldc.upenn.edu/LDC2026S08">CALLHOME American =
English Second Edition</a><i><br>
</i><a href=3D"https://catalog.ldc.upenn.edu/LDC2026L05">CALLHOME American =
English Lexicon (PRONLEX) Second Edition</a><br>
<br>
<b><o:p></o:p></b></span></p>
<div class=3D"MsoNormal" align=3D"center" style=3D"text-align:center"><a na=
me=3D"_Hlk215751265"></a><a name=3D"_Hlk215750926"><span style=3D"mso-bookm=
ark:_Hlk215751265"><b><span style=3D"font-size:11.0pt">
<hr size=3D"2" width=3D"100%" noshade=3D"" style=3D"color:black" align=3D"c=
enter">
</span></b></span></a></div>
<p class=3D"MsoNormal"><span style=3D"mso-bookmark:_Hlk215750926"><span sty=
le=3D"mso-bookmark:_Hlk215751265"><b><span style=3D"font-size:11.0pt">Fall =
2026 LDC data scholarship program&nbsp;<br>
</span></b></span></span><span style=3D"mso-bookmark:_Hlk215750926"><span s=
tyle=3D"mso-bookmark:_Hlk215751265"><span style=3D"font-size:11.0pt">Studen=
t applications for the Fall 2026 LDC data scholarship program are being acc=
epted now through September 15, 2026.
 This program provides&nbsp;eligible&nbsp;students with&nbsp;no-cost&nbsp;a=
ccess to LDC data.&nbsp;Students must&nbsp;complete&nbsp;an application con=
sisting of a data use proposal and letter of support from their advisor. Fo=
r application requirements and program rules, visit the&nbsp;</span></span>=
</span><a href=3D"https://www.ldc.upenn.edu/language-resources/data/data-sc=
holarships"><span style=3D"mso-bookmark:_Hlk215750926"><span style=3D"mso-b=
ookmark:_Hlk215751265"><span style=3D"font-size:11.0pt">LDC
 Data Scholarships</span></span></span><span style=3D"mso-bookmark:_Hlk2157=
50926"><span style=3D"mso-bookmark:_Hlk215751265"></span></span></a><span s=
tyle=3D"mso-bookmark:_Hlk215750926"><span style=3D"mso-bookmark:_Hlk2157512=
65"><span style=3D"font-size:11.0pt"> page.<b>&nbsp;</b></span></span></spa=
n><b><span style=3D"font-size:11.0pt"><o:p></o:p></span></b></p>
<div class=3D"MsoNormal" align=3D"center" style=3D"text-align:center"><b><s=
pan style=3D"font-size:11.0pt">
<hr size=3D"2" width=3D"100%" noshade=3D"" style=3D"color:black" align=3D"c=
enter">
</span></b></div>
<p class=3D"MsoNormal"><i><span style=3D"font-size:11.0pt"><br>
New publications:<a name=3D"_Hlk100573448"></a><a name=3D"_Hlk97739035"></a=
><a name=3D"_Hlk97739153"></a><a name=3D"_Hlk216348385"><span style=3D"mso-=
bookmark:_Hlk97739153"><span style=3D"mso-bookmark:_Hlk97739035"><span styl=
e=3D"mso-bookmark:_Hlk100573448"><br>
</span></span></span></a></span></i><span style=3D"mso-bookmark:_Hlk2163483=
85"><span style=3D"mso-bookmark:_Hlk97739153"><span style=3D"mso-bookmark:_=
Hlk97739035"><span style=3D"mso-bookmark:_Hlk100573448"></span></span></spa=
n></span><a href=3D"https://catalog.ldc.upenn.edu/LDC2026S09"><span style=
=3D"mso-bookmark:_Hlk216348385"><span style=3D"mso-bookmark:_Hlk97739153"><=
span style=3D"mso-bookmark:_Hlk97739035"><span style=3D"mso-bookmark:_Hlk10=
0573448"><span style=3D"font-size:11.0pt">2012
 NIST Speaker Recognition Evaluation Test Set</span></span></span></span></=
span><span style=3D"mso-bookmark:_Hlk216348385"><span style=3D"mso-bookmark=
:_Hlk97739153"><span style=3D"mso-bookmark:_Hlk97739035"><span style=3D"mso=
-bookmark:_Hlk100573448"></span></span></span></span></a><span style=3D"mso=
-bookmark:_Hlk216348385"><span style=3D"mso-bookmark:_Hlk97739153"><span st=
yle=3D"mso-bookmark:_Hlk97739035"><span style=3D"mso-bookmark:_Hlk100573448=
"><span style=3D"font-size:11.0pt">
 was developed by LDC and NIST and contains 10,321 hours of English convers=
ational telephone speech and in-person recorded studio sessions for evaluat=
ion and modeling, along with answer keys, trial files, and documentation fr=
om the NIST-sponsored
</span></span></span></span></span><a href=3D"https://www.nist.gov/itl/iad/=
mltg/speaker-recognition-evaluation-2012"><span style=3D"mso-bookmark:_Hlk2=
16348385"><span style=3D"mso-bookmark:_Hlk97739153"><span style=3D"mso-book=
mark:_Hlk97739035"><span style=3D"mso-bookmark:_Hlk100573448"><span style=
=3D"font-size:11.0pt">2012
 Speaker Recognition Evaluation (SRE12)</span></span></span></span></span><=
span style=3D"mso-bookmark:_Hlk216348385"><span style=3D"mso-bookmark:_Hlk9=
7739153"><span style=3D"mso-bookmark:_Hlk97739035"><span style=3D"mso-bookm=
ark:_Hlk100573448"></span></span></span></span></a><span style=3D"mso-bookm=
ark:_Hlk216348385"><span style=3D"mso-bookmark:_Hlk97739153"><span style=3D=
"mso-bookmark:_Hlk97739035"><span style=3D"mso-bookmark:_Hlk100573448"><spa=
n style=3D"font-size:11.0pt">.
 SRE12 introduced a revised evaluation structure in which training data for=
 target speakers was drawn from prior SRE corpora developed by LDC and was =
provided in advance of the evaluation period.
<i><br>
<br>
</i>Test data was drawn from Mixer 7 English Speech </span></span></span></=
span></span><a href=3D"https://catalog.ldc.upenn.edu/LDC2025S08"><span styl=
e=3D"mso-bookmark:_Hlk216348385"><span style=3D"mso-bookmark:_Hlk97739153">=
<span style=3D"mso-bookmark:_Hlk97739035"><span style=3D"mso-bookmark:_Hlk1=
00573448"><span style=3D"font-size:11.0pt">(LDC2025S08)</span></span></span=
></span></span><span style=3D"mso-bookmark:_Hlk216348385"><span style=3D"ms=
o-bookmark:_Hlk97739153"><span style=3D"mso-bookmark:_Hlk97739035"><span st=
yle=3D"mso-bookmark:_Hlk100573448"></span></span></span></span></a><span st=
yle=3D"mso-bookmark:_Hlk216348385"><span style=3D"mso-bookmark:_Hlk97739153=
"><span style=3D"mso-bookmark:_Hlk97739035"><span style=3D"mso-bookmark:_Hl=
k100573448"><span style=3D"font-size:11.0pt">
 and REMIX Telephone Collection </span></span></span></span></span><a href=
=3D"https://catalog.ldc.upenn.edu/LDC2023S09"><span style=3D"mso-bookmark:_=
Hlk216348385"><span style=3D"mso-bookmark:_Hlk97739153"><span style=3D"mso-=
bookmark:_Hlk97739035"><span style=3D"mso-bookmark:_Hlk100573448"><span sty=
le=3D"font-size:11.0pt">(LDC2023S09)</span></span></span></span></span><spa=
n style=3D"mso-bookmark:_Hlk216348385"><span style=3D"mso-bookmark:_Hlk9773=
9153"><span style=3D"mso-bookmark:_Hlk97739035"><span style=3D"mso-bookmark=
:_Hlk100573448"></span></span></span></span></a><span style=3D"mso-bookmark=
:_Hlk216348385"><span style=3D"mso-bookmark:_Hlk97739153"><span style=3D"ms=
o-bookmark:_Hlk97739035"><span style=3D"mso-bookmark:_Hlk100573448"><span s=
tyle=3D"font-size:11.0pt">.
 Those datasets also provided segments for modeling data; other modeling se=
gments were drawn from Mixer 3 Speech
</span></span></span></span></span><a href=3D"https://catalog.ldc.upenn.edu=
/LDC2023S02"><span style=3D"mso-bookmark:_Hlk216348385"><span style=3D"mso-=
bookmark:_Hlk97739153"><span style=3D"mso-bookmark:_Hlk97739035"><span styl=
e=3D"mso-bookmark:_Hlk100573448"><span style=3D"font-size:11.0pt">(LDC2023S=
02)</span></span></span></span></span><span style=3D"mso-bookmark:_Hlk21634=
8385"><span style=3D"mso-bookmark:_Hlk97739153"><span style=3D"mso-bookmark=
:_Hlk97739035"><span style=3D"mso-bookmark:_Hlk100573448"></span></span></s=
pan></span></a><span style=3D"mso-bookmark:_Hlk216348385"><span style=3D"ms=
o-bookmark:_Hlk97739153"><span style=3D"mso-bookmark:_Hlk97739035"><span st=
yle=3D"mso-bookmark:_Hlk100573448"><span style=3D"font-size:11.0pt">,
 Mixer 4 and 5 Speech </span></span></span></span></span><a href=3D"https:/=
/catalog.ldc.upenn.edu/LDC2020S03"><span style=3D"mso-bookmark:_Hlk21634838=
5"><span style=3D"mso-bookmark:_Hlk97739153"><span style=3D"mso-bookmark:_H=
lk97739035"><span style=3D"mso-bookmark:_Hlk100573448"><span style=3D"font-=
size:11.0pt">(LDC2020S03)</span></span></span></span></span><span style=3D"=
mso-bookmark:_Hlk216348385"><span style=3D"mso-bookmark:_Hlk97739153"><span=
 style=3D"mso-bookmark:_Hlk97739035"><span style=3D"mso-bookmark:_Hlk100573=
448"></span></span></span></span></a><span style=3D"mso-bookmark:_Hlk216348=
385"><span style=3D"mso-bookmark:_Hlk97739153"><span style=3D"mso-bookmark:=
_Hlk97739035"><span style=3D"mso-bookmark:_Hlk100573448"><span style=3D"fon=
t-size:11.0pt">,
 and Mixer 6 Speech </span></span></span></span></span><a href=3D"https://c=
atalog.ldc.upenn.edu/LDC2013S03"><span style=3D"mso-bookmark:_Hlk216348385"=
><span style=3D"mso-bookmark:_Hlk97739153"><span style=3D"mso-bookmark:_Hlk=
97739035"><span style=3D"mso-bookmark:_Hlk100573448"><span style=3D"font-si=
ze:11.0pt">(LDC2013S03)</span></span></span></span></span><span style=3D"ms=
o-bookmark:_Hlk216348385"><span style=3D"mso-bookmark:_Hlk97739153"><span s=
tyle=3D"mso-bookmark:_Hlk97739035"><span style=3D"mso-bookmark:_Hlk10057344=
8"></span></span></span></span></a><span style=3D"mso-bookmark:_Hlk21634838=
5"><span style=3D"mso-bookmark:_Hlk97739153"><span style=3D"mso-bookmark:_H=
lk97739035"><span style=3D"mso-bookmark:_Hlk100573448"><span style=3D"font-=
size:11.0pt">.
 The test data contains English speech only; some non-English speech is con=
tained in modeling segments.<i><br>
<br>
</i>This release is comprised of 130,844 test segments, specifically, 83,77=
8 call segments and 47,066 interview segments. Modeling data consists of 46=
,948 segments.<i><br>
<br>
</i>2026 members can access this corpus through their LDC accounts. Non-mem=
bers may license this data for a fee.<br>
<br>
<b><o:p></o:p></b></span></span></span></span></span></p>
<span style=3D"mso-bookmark:_Hlk97739153"></span><span style=3D"mso-bookmar=
k:_Hlk216348385"></span>
<p class=3D"MsoNormal"><span style=3D"mso-bookmark:_Hlk97739035"><span styl=
e=3D"mso-bookmark:_Hlk100573448"><a name=3D"_Hlk174374006"><span style=3D"f=
ont-size:11.0pt">*</span></a><a name=3D"_Hlk119070711"><span style=3D"mso-b=
ookmark:_Hlk174374006"><br>
<br>
</span></a></span></span><span style=3D"mso-bookmark:_Hlk97739035"><span st=
yle=3D"mso-bookmark:_Hlk100573448"><span style=3D"mso-bookmark:_Hlk11907071=
1"><span style=3D"mso-bookmark:_Hlk174374006"><span style=3D"font-size:11.0=
pt"><o:p></o:p></span></span></span></span></span></p>
<p class=3D"MsoNormal"><span style=3D"mso-bookmark:_Hlk97739035"><span styl=
e=3D"mso-bookmark:_Hlk100573448"><span style=3D"mso-bookmark:_Hlk119070711"=
><span style=3D"mso-bookmark:_Hlk174374006"></span></span></span></span><a =
href=3D"https://catalog.ldc.upenn.edu/LDC2026S08"><span style=3D"mso-bookma=
rk:_Hlk97739035"><span style=3D"mso-bookmark:_Hlk100573448"><span style=3D"=
mso-bookmark:_Hlk119070711"><span style=3D"mso-bookmark:_Hlk174374006"><spa=
n style=3D"font-size:11.0pt">CALLHOME
 American English Second Edition</span></span></span></span></span><span st=
yle=3D"mso-bookmark:_Hlk97739035"><span style=3D"mso-bookmark:_Hlk100573448=
"><span style=3D"mso-bookmark:_Hlk119070711"><span style=3D"mso-bookmark:_H=
lk174374006"></span></span></span></span></a><span style=3D"mso-bookmark:_H=
lk97739035"><span style=3D"mso-bookmark:_Hlk100573448"><span style=3D"mso-b=
ookmark:_Hlk119070711"><span style=3D"mso-bookmark:_Hlk174374006"><span sty=
le=3D"font-size:11.0pt">&nbsp;was
 developed by LDC and contains 56 hours of speech from 120 unscripted telep=
hone conversations between native American English speakers. This publicati=
on is a re-release of the original CALLHOME American English collection, co=
mbining CALLHOME American English
 Speech </span></span></span></span></span><a href=3D"https://catalog.ldc.u=
penn.edu/LDC97S42"><span style=3D"mso-bookmark:_Hlk97739035"><span style=3D=
"mso-bookmark:_Hlk100573448"><span style=3D"mso-bookmark:_Hlk119070711"><sp=
an style=3D"mso-bookmark:_Hlk174374006"><span style=3D"font-size:11.0pt">(L=
DC97S42)</span></span></span></span></span><span style=3D"mso-bookmark:_Hlk=
97739035"><span style=3D"mso-bookmark:_Hlk100573448"><span style=3D"mso-boo=
kmark:_Hlk119070711"><span style=3D"mso-bookmark:_Hlk174374006"></span></sp=
an></span></span></a><span style=3D"mso-bookmark:_Hlk97739035"><span style=
=3D"mso-bookmark:_Hlk100573448"><span style=3D"mso-bookmark:_Hlk119070711">=
<span style=3D"mso-bookmark:_Hlk174374006"><span style=3D"font-size:11.0pt"=
>
 and CALLHOME American English Transcripts </span></span></span></span></sp=
an><a href=3D"https://catalog.ldc.upenn.edu/LDC97T14"><span style=3D"mso-bo=
okmark:_Hlk97739035"><span style=3D"mso-bookmark:_Hlk100573448"><span style=
=3D"mso-bookmark:_Hlk119070711"><span style=3D"mso-bookmark:_Hlk174374006">=
<span style=3D"font-size:11.0pt">(LDC97T14)</span></span></span></span></sp=
an><span style=3D"mso-bookmark:_Hlk97739035"><span style=3D"mso-bookmark:_H=
lk100573448"><span style=3D"mso-bookmark:_Hlk119070711"><span style=3D"mso-=
bookmark:_Hlk174374006"></span></span></span></span></a><span style=3D"mso-=
bookmark:_Hlk97739035"><span style=3D"mso-bookmark:_Hlk100573448"><span sty=
le=3D"mso-bookmark:_Hlk119070711"><span style=3D"mso-bookmark:_Hlk174374006=
"><span style=3D"font-size:11.0pt">,
 with additional transcription and updated directory structure, file format=
s, and documentation.<br>
<br>
This release contains the 120 telephone conversations published in CALLHOME=
 American English Speech which represented training and development data an=
d a subset of evaluation data. Participants spoke on topics of their choice=
 in a single telephone call lasting
 up to 30 minutes. Calls were manually audited for gender, language, record=
ing quality, channel characteristics, dialect, and region. For this second =
edition, all audio was converted from SPHERE files to FLAC format, and the =
original training/development/evaluation
 partitioning was removed. <br>
<br>
This release also features revised transcripts conforming to updated LDC tr=
anscription guidelines that addressed normalization of annotation formats, =
standardization of speaker-produced and background noises, application of f=
oreign-language marking, whitespace
 cleanup, and corrections and consistency fixes.<br>
<br>
The CALLHOME series consists of telephone conversations and transcripts dev=
eloped by LDC and Rutgers, The State University of New Jersey, in support o=
f research in speaker identification, language identification, and related =
technologies. Languages in the series
 include American English, Egyptian Arabic, German, Japanese, Mandarin Chin=
ese, and Spanish.<br>
<br>
2026 members can access this corpus through their LDC accounts. Non-members=
 may license this data for a fee.<br>
<br>
<o:p></o:p></span></span></span></span></span></p>
<p class=3D"MsoNormal"><span style=3D"mso-bookmark:_Hlk97739035"><span styl=
e=3D"mso-bookmark:_Hlk100573448"><span style=3D"mso-bookmark:_Hlk119070711"=
><span style=3D"mso-bookmark:_Hlk174374006"><a name=3D"_Hlk219206384"></a><=
a name=3D"_Hlk174372776"><span style=3D"mso-bookmark:_Hlk219206384"><span s=
tyle=3D"font-size:11.0pt">*<br>
<br>
<o:p></o:p></span></span></a></span></span></span></span></p>
<p class=3D"MsoNormal"><span style=3D"mso-bookmark:_Hlk97739035"><span styl=
e=3D"mso-bookmark:_Hlk100573448"><span style=3D"mso-bookmark:_Hlk119070711"=
><span style=3D"mso-bookmark:_Hlk174374006"><span style=3D"mso-bookmark:_Hl=
k174372776"><span style=3D"mso-bookmark:_Hlk219206384"></span></span></span=
></span></span></span><a href=3D"https://catalog.ldc.upenn.edu/LDC2026L05">=
<span style=3D"mso-bookmark:_Hlk97739035"><span style=3D"mso-bookmark:_Hlk1=
00573448"><span style=3D"mso-bookmark:_Hlk119070711"><span style=3D"mso-boo=
kmark:_Hlk174374006"><span style=3D"mso-bookmark:_Hlk174372776"><span style=
=3D"mso-bookmark:_Hlk219206384"><span style=3D"font-size:11.0pt">CALLHOME
 American English Lexicon (PRONLEX) Second Edition</span></span></span></sp=
an></span></span></span><span style=3D"mso-bookmark:_Hlk97739035"><span sty=
le=3D"mso-bookmark:_Hlk100573448"><span style=3D"mso-bookmark:_Hlk119070711=
"><span style=3D"mso-bookmark:_Hlk174374006"><span style=3D"mso-bookmark:_H=
lk174372776"><span style=3D"mso-bookmark:_Hlk219206384"></span></span></spa=
n></span></span></span></a><span style=3D"mso-bookmark:_Hlk97739035"><span =
style=3D"mso-bookmark:_Hlk100573448"><span style=3D"mso-bookmark:_Hlk119070=
711"><span style=3D"mso-bookmark:_Hlk174374006"><span style=3D"mso-bookmark=
:_Hlk174372776"><span style=3D"mso-bookmark:_Hlk219206384"><span style=3D"f=
ont-size:11.0pt">
 was developed by LDC and contains 90,988 English words with citation-form =
pronunciations. This second edition updates file formats, directory structu=
re, and documentation. The first edition is available as CALLHOME American =
English Lexicon (PRONLEX)
</span></span></span></span></span></span></span><a href=3D"https://catalog=
.ldc.upenn.edu/LDC97L20"><span style=3D"mso-bookmark:_Hlk97739035"><span st=
yle=3D"mso-bookmark:_Hlk100573448"><span style=3D"mso-bookmark:_Hlk11907071=
1"><span style=3D"mso-bookmark:_Hlk174374006"><span style=3D"mso-bookmark:_=
Hlk174372776"><span style=3D"mso-bookmark:_Hlk219206384"><span style=3D"fon=
t-size:11.0pt">(LDC97L20)</span></span></span></span></span></span></span><=
span style=3D"mso-bookmark:_Hlk97739035"><span style=3D"mso-bookmark:_Hlk10=
0573448"><span style=3D"mso-bookmark:_Hlk119070711"><span style=3D"mso-book=
mark:_Hlk174374006"><span style=3D"mso-bookmark:_Hlk174372776"><span style=
=3D"mso-bookmark:_Hlk219206384"></span></span></span></span></span></span><=
/a><span style=3D"mso-bookmark:_Hlk97739035"><span style=3D"mso-bookmark:_H=
lk100573448"><span style=3D"mso-bookmark:_Hlk119070711"><span style=3D"mso-=
bookmark:_Hlk174374006"><span style=3D"mso-bookmark:_Hlk174372776"><span st=
yle=3D"mso-bookmark:_Hlk219206384"><span style=3D"font-size:11.0pt">.
<br>
<br>
The words in the lexicon were derived from Wall Street Journal text used in=
 the continuous speech recognition publication series CSR-1 WSJ0 Complete (=
</span></span></span></span></span></span></span><a href=3D"https://catalog=
.ldc.upenn.edu/LDC93S6A"><span style=3D"mso-bookmark:_Hlk97739035"><span st=
yle=3D"mso-bookmark:_Hlk100573448"><span style=3D"mso-bookmark:_Hlk11907071=
1"><span style=3D"mso-bookmark:_Hlk174374006"><span style=3D"mso-bookmark:_=
Hlk174372776"><span style=3D"mso-bookmark:_Hlk219206384"><span style=3D"fon=
t-size:11.0pt">LDC93S6A</span></span></span></span></span></span></span><sp=
an style=3D"mso-bookmark:_Hlk97739035"><span style=3D"mso-bookmark:_Hlk1005=
73448"><span style=3D"mso-bookmark:_Hlk119070711"><span style=3D"mso-bookma=
rk:_Hlk174374006"><span style=3D"mso-bookmark:_Hlk174372776"><span style=3D=
"mso-bookmark:_Hlk219206384"></span></span></span></span></span></span></a>=
<span style=3D"mso-bookmark:_Hlk97739035"><span style=3D"mso-bookmark:_Hlk1=
00573448"><span style=3D"mso-bookmark:_Hlk119070711"><span style=3D"mso-boo=
kmark:_Hlk174374006"><span style=3D"mso-bookmark:_Hlk174372776"><span style=
=3D"mso-bookmark:_Hlk219206384"><span style=3D"font-size:11.0pt">),
</span></span></span></span></span></span></span><a href=3D"https://isip.pi=
conepress.com/projects/switchboard/"><span style=3D"mso-bookmark:_Hlk977390=
35"><span style=3D"mso-bookmark:_Hlk100573448"><span style=3D"mso-bookmark:=
_Hlk119070711"><span style=3D"mso-bookmark:_Hlk174374006"><span style=3D"ms=
o-bookmark:_Hlk174372776"><span style=3D"mso-bookmark:_Hlk219206384"><span =
style=3D"font-size:11.0pt">transcripts</span></span></span></span></span></=
span></span><span style=3D"mso-bookmark:_Hlk97739035"><span style=3D"mso-bo=
okmark:_Hlk100573448"><span style=3D"mso-bookmark:_Hlk119070711"><span styl=
e=3D"mso-bookmark:_Hlk174374006"><span style=3D"mso-bookmark:_Hlk174372776"=
><span style=3D"mso-bookmark:_Hlk219206384"></span></span></span></span></s=
pan></span></a><span style=3D"mso-bookmark:_Hlk97739035"><span style=3D"mso=
-bookmark:_Hlk100573448"><span style=3D"mso-bookmark:_Hlk119070711"><span s=
tyle=3D"mso-bookmark:_Hlk174374006"><span style=3D"mso-bookmark:_Hlk1743727=
76"><span style=3D"mso-bookmark:_Hlk219206384"><span style=3D"font-size:11.=
0pt">
 from the Switchboard telephone collection </span></span></span></span></sp=
an></span></span><a href=3D"https://catalog.ldc.upenn.edu/LDC97S62"><span s=
tyle=3D"mso-bookmark:_Hlk97739035"><span style=3D"mso-bookmark:_Hlk10057344=
8"><span style=3D"mso-bookmark:_Hlk119070711"><span style=3D"mso-bookmark:_=
Hlk174374006"><span style=3D"mso-bookmark:_Hlk174372776"><span style=3D"mso=
-bookmark:_Hlk219206384"><span style=3D"font-size:11.0pt">(LDC97S62)</span>=
</span></span></span></span></span></span><span style=3D"mso-bookmark:_Hlk9=
7739035"><span style=3D"mso-bookmark:_Hlk100573448"><span style=3D"mso-book=
mark:_Hlk119070711"><span style=3D"mso-bookmark:_Hlk174374006"><span style=
=3D"mso-bookmark:_Hlk174372776"><span style=3D"mso-bookmark:_Hlk219206384">=
</span></span></span></span></span></span></a><span style=3D"mso-bookmark:_=
Hlk97739035"><span style=3D"mso-bookmark:_Hlk100573448"><span style=3D"mso-=
bookmark:_Hlk119070711"><span style=3D"mso-bookmark:_Hlk174374006"><span st=
yle=3D"mso-bookmark:_Hlk174372776"><span style=3D"mso-bookmark:_Hlk21920638=
4"><span style=3D"font-size:11.0pt">,
 and transcripts representing unscripted telephone conversations between na=
tive American English speakers contained in CALLHOME American English Secon=
d Edition
</span></span></span></span></span></span></span><a href=3D"https://catalog=
.ldc.upenn.edu/LDC2026S08"><span style=3D"mso-bookmark:_Hlk97739035"><span =
style=3D"mso-bookmark:_Hlk100573448"><span style=3D"mso-bookmark:_Hlk119070=
711"><span style=3D"mso-bookmark:_Hlk174374006"><span style=3D"mso-bookmark=
:_Hlk174372776"><span style=3D"mso-bookmark:_Hlk219206384"><span style=3D"f=
ont-size:11.0pt">(LDC2026S08)</span></span></span></span></span></span></sp=
an><span style=3D"mso-bookmark:_Hlk97739035"><span style=3D"mso-bookmark:_H=
lk100573448"><span style=3D"mso-bookmark:_Hlk119070711"><span style=3D"mso-=
bookmark:_Hlk174374006"><span style=3D"mso-bookmark:_Hlk174372776"><span st=
yle=3D"mso-bookmark:_Hlk219206384"></span></span></span></span></span></spa=
n></a><span style=3D"mso-bookmark:_Hlk97739035"><span style=3D"mso-bookmark=
:_Hlk100573448"><span style=3D"mso-bookmark:_Hlk119070711"><span style=3D"m=
so-bookmark:_Hlk174374006"><span style=3D"mso-bookmark:_Hlk174372776"><span=
 style=3D"mso-bookmark:_Hlk219206384"><span style=3D"font-size:11.0pt">.<br=
>
<br>
PRONLEX transcription is a phonemic transcription system designed to suppor=
t speech recognition by providing a consistent and simplified representatio=
n of how words are pronounced in standard American English that allows vari=
ation to be generated later to avoid
 listing many pronunciation variations for each word. This single systemati=
c base form can be expanded through rules or modeling. The transcription wa=
s created using a modified
</span></span></span></span></span></span></span><a href=3D"https://learniu=
s.com/slp/3+Speech+Production%2C+Perception+and+Phonetics/4+Phonetics/2+Pho=
netic+Alphabet/ARPAbet"><span style=3D"mso-bookmark:_Hlk97739035"><span sty=
le=3D"mso-bookmark:_Hlk100573448"><span style=3D"mso-bookmark:_Hlk119070711=
"><span style=3D"mso-bookmark:_Hlk174374006"><span style=3D"mso-bookmark:_H=
lk174372776"><span style=3D"mso-bookmark:_Hlk219206384"><span style=3D"font=
-size:11.0pt">ARPABET
 phoneme set</span></span></span></span></span></span></span><span style=3D=
"mso-bookmark:_Hlk97739035"><span style=3D"mso-bookmark:_Hlk100573448"><spa=
n style=3D"mso-bookmark:_Hlk119070711"><span style=3D"mso-bookmark:_Hlk1743=
74006"><span style=3D"mso-bookmark:_Hlk174372776"><span style=3D"mso-bookma=
rk:_Hlk219206384"></span></span></span></span></span></span></a><span style=
=3D"mso-bookmark:_Hlk97739035"><span style=3D"mso-bookmark:_Hlk100573448"><=
span style=3D"mso-bookmark:_Hlk119070711"><span style=3D"mso-bookmark:_Hlk1=
74374006"><span style=3D"mso-bookmark:_Hlk174372776"><span style=3D"mso-boo=
kmark:_Hlk219206384"><span style=3D"font-size:11.0pt">.<br>
<br>
The lexicon contains three tab-separated information fields: (1) word: orth=
ographic representation of word; (2) pron: transcribed citation-form pronun=
ciations using modified ARPABET phoneme set; and (3) comments: (OPTIONAL) c=
omment on the entry. It is presented
 as a tab-delimited TSV file encoded in UTF-8 format and includes a pronunc=
iation dictionary derived from the lexicon in UTF-8 encoded
</span></span></span></span></span></span></span><a href=3D"https://stdlib.=
io/docs/api/latest/@stdlib/datasets/cmudict"><span style=3D"mso-bookmark:_H=
lk97739035"><span style=3D"mso-bookmark:_Hlk100573448"><span style=3D"mso-b=
ookmark:_Hlk119070711"><span style=3D"mso-bookmark:_Hlk174374006"><span sty=
le=3D"mso-bookmark:_Hlk174372776"><span style=3D"mso-bookmark:_Hlk219206384=
"><span style=3D"font-size:11.0pt">CMUdict</span></span></span></span></spa=
n></span></span><span style=3D"mso-bookmark:_Hlk97739035"><span style=3D"ms=
o-bookmark:_Hlk100573448"><span style=3D"mso-bookmark:_Hlk119070711"><span =
style=3D"mso-bookmark:_Hlk174374006"><span style=3D"mso-bookmark:_Hlk174372=
776"><span style=3D"mso-bookmark:_Hlk219206384"></span></span></span></span=
></span></span></a><span style=3D"mso-bookmark:_Hlk97739035"><span style=3D=
"mso-bookmark:_Hlk100573448"><span style=3D"mso-bookmark:_Hlk119070711"><sp=
an style=3D"mso-bookmark:_Hlk174374006"><span style=3D"mso-bookmark:_Hlk174=
372776"><span style=3D"mso-bookmark:_Hlk219206384"><span style=3D"font-size=
:11.0pt">
 format. <br>
<br>
2026 members can access this corpus through their LDC accounts provided the=
y have submitted a completed copy of the special license agreement. Non-mem=
bers may license this data for a fee.</span></span></span></span></span></s=
pan></span><span style=3D"mso-bookmark:_Hlk97739035"><span style=3D"mso-boo=
kmark:_Hlk100573448"><span style=3D"font-size:11.0pt"><br>
<br>
<b>To unsubscribe from this newsletter, log in to your </b></span></span></=
span><a href=3D"https://catalog.ldc.upenn.edu/login"><span style=3D"mso-boo=
kmark:_Hlk97739035"><span style=3D"mso-bookmark:_Hlk100573448"><b><span sty=
le=3D"font-size:11.0pt">LDC account</span></b></span></span><span style=3D"=
mso-bookmark:_Hlk97739035"><span style=3D"mso-bookmark:_Hlk100573448"></spa=
n></span></a><span style=3D"mso-bookmark:_Hlk97739035"><span style=3D"mso-b=
ookmark:_Hlk100573448"><b><span style=3D"font-size:11.0pt">
 and uncheck the box next to &#8220;Receive Newsletter&#8221; under Account=
 Options or contact LDC for assistance.
</span></b></span></span><span style=3D"mso-bookmark:_Hlk100573448"></span>=
<span style=3D"mso-bookmark:_Hlk97739035"></span><span style=3D"font-size:1=
1.0pt"><o:p></o:p></span></p>
<p class=3D"MsoNormal"><span style=3D"font-size:11.0pt"><o:p>&nbsp;</o:p></=
span></p>
<p class=3D"MsoNormal"><span style=3D"font-size:11.0pt">Membership Coordina=
tor<o:p></o:p></span></p>
<p class=3D"MsoNormal"><span style=3D"font-size:11.0pt"><a href=3D"ldc.upen=
n.edu">Linguistic Data Consortium</a><o:p></o:p></span></p>
<p class=3D"MsoNormal"><span style=3D"font-size:11.0pt">University of Penns=
ylvania<o:p></o:p></span></p>
<p class=3D"MsoNormal"><span style=3D"font-size:11.0pt">T: +1-215-573-1275<=
o:p></o:p></span></p>
<p class=3D"MsoNormal"><span style=3D"font-size:11.0pt">E: <a href=3D"mailt=
o:[email protected]">
[email protected]</a><o:p></o:p></span></p>
<p class=3D"MsoNormal"><span style=3D"font-size:11.0pt">M: 3600 Market St. =
Suite 810<o:p></o:p></span></p>
<p class=3D"MsoNormal"><span style=3D"font-size:11.0pt">&nbsp;&nbsp;&nbsp;&=
nbsp;&nbsp; Philadelphia, PA 19104 <o:p>
</o:p></span></p>
<p class=3D"MsoNormal"><span style=3D"font-size:11.0pt"><o:p>&nbsp;</o:p></=
span></p>
<p class=3D"MsoNormal"><span style=3D"font-size:11.0pt"><o:p>&nbsp;</o:p></=
span></p>
<p class=3D"MsoNormal"><o:p>&nbsp;</o:p></p>
</div>
</body>
</html>

--_000_BN0PR10MB50295919A09C53CF0FC6BCB3F9F82BN0PR10MB5029namp_--

--===============5788192565136502589==
Content-Type: text/plain; charset="us-ascii"
MIME-Version: 1.0
Content-Transfer-Encoding: 7bit
Content-Disposition: inline

_______________________________________________
Corpora mailing list -- [email protected]
https://list.elra.info/mailman3/postorius/lists/corpora.list.elra.info/
To unsubscribe send an email to [email protected]

--===============5788192565136502589==--