Words' list relating to word containing wildcard *, ?, #
eric leroy <[email protected]> Fri, 6 Nov 2020 10:15:54 +0100
| Newsgroups | gmane.comp.search.snowball |
|---|---|
| Message-ID | <CAFs9Jx17G7TgN9=uGf050MaaaipYyR5y-mphyHWUfJRgMVBEyg@mail.gmail.com> |
--===============7916966098435099961== Content-Type: multipart/alternative; boundary="00000000000008064605b36cab92" --00000000000008064605b36cab92 Content-Type: text/plain; charset="UTF-8" Content-Transfer-Encoding: quoted-printable Hi everybody, I'm new here ! Here's my topic, and thank you all for your help and advice.As said in the subject, I'd like to obtain a words' list relating to words containing wildcard *, ?, #. The reason is that I=E2=80=99m trying to migrate a dictionary from a platfo= rm that allowed me using wildcard *, ?, # associates with part of words (as a single entry in my dictionary or in word sequences) to another platform that doesn=E2=80=99t take into account such characters and force me to crea= te a single line for each declination. Using a snowball =C2=AB * =C2=BB allowed = me in my present dictionary, to capture all part of texts relating to these variations (pluriel, gender, grammatical declinaison, etc.). For example, SUPPORT* will mean SUPPORT, SUPPORTS, SUPPORTING, SUPPORTIVE, SUPPORTER, etc. While the following word pattern: *SUPPORT* will also substitute all words with the substring "SUPPORT" in it, such as UNSUPPORTEDLY, UNSUPPORTED, etc= . An expression that includes several words may also be substituted by joining the various words with underline characters. For example, the expression "going out" GO*_OUT. But my needs go beyond the snowball as wildcards such as *, ?, # are supported in my dictionary: =C2=AB ? =C2=BB to replace any character in a = word, =C2=AB # =C2=BB to replace any number, etc. Therefore, I need to migrate my actual dictionary (French words) that contains thousands of rows with ITEMS containing wildcards: is there a solution that could allow me to give, for each such word, all the corresponding words? Thanks a lot for your suggestions, Eric --00000000000008064605b36cab92 Content-Type: text/html; charset="UTF-8" Content-Transfer-Encoding: quoted-printable <div dir=3D"ltr"><div><div dir=3D"ltr" class=3D"gmail_signature" data-smart= mail=3D"gmail_signature"><div dir=3D"ltr"><div><div dir=3D"ltr"><div dir=3D= "ltr"><div dir=3D"ltr"><div dir=3D"ltr"><div dir=3D"ltr"><div dir=3D"ltr"><= div dir=3D"ltr"><div dir=3D"ltr"><div dir=3D"ltr"><div dir=3D"ltr"><p style= =3D"font-size:small">Hi everybody,<br>I'm new here !<br>Here's my t= opic, and thank you all for your help and advice.As said in the subject, I&= #39;d like to obtain a words' list relating to words containing wildcar= d *, ?, #.<br>The reason is that I=E2=80=99m trying to migrate a dictionary= from a platform that allowed me using wildcard *, ?, # associates with par= t of words (as a single entry in my dictionary or in word sequences) to ano= ther platform that doesn=E2=80=99t take into account such characters and fo= rce me to create a single line for each declination. Using a snowball =C2= =AB * =C2=BB allowed me in my present dictionary, to capture all part of te= xts relating to these variations (pluriel, gender, grammatical declinaison,= etc.).<br><br>For example, SUPPORT* will mean SUPPORT, SUPPORTS, SUPPORTIN= G, SUPPORTIVE, SUPPORTER, etc.<br>While the following word pattern: *SUPPOR= T* will also substitute all words with the substring "SUPPORT" in= it, such as UNSUPPORTEDLY, UNSUPPORTED, etc.<br>An expression that include= s several words may also be substituted by joining the various words with u= nderline characters. For example, the expression "going out" =C2= =A0GO*_OUT.<br>=C2=A0<br>But my needs go beyond the snowball as wildcards s= uch as =C2=A0*, =C2=A0?, # are supported in my dictionary: =C2=A0=C2=AB ? = =C2=BB to replace any character in a word, =C2=AB # =C2=BB to replace any n= umber, etc.<br><br>Therefore, I need to migrate my actual dictionary (Frenc= h words) that contains thousands of rows with ITEMS containing wildcards: = =C2=A0is there a solution that could allow me to give, for each such word, = all the corresponding words?<br><br>Thanks a lot for your suggestions,<br><= br>Eric<br></p></div></div></div></div></div></div></div></div></div></div>= </div></div></div></div></div> --00000000000008064605b36cab92-- --===============7916966098435099961== Content-Type: text/plain; charset="us-ascii" MIME-Version: 1.0 Content-Transfer-Encoding: 7bit Content-Disposition: inline _______________________________________________ Snowball-discuss mailing list [email protected] https://lists.tartarus.org/mailman/listinfo/snowball-discuss --===============7916966098435099961==--