Re: Problem of large data sets
Shu-Ju Tu <[email protected]> Thu, 18 Apr 2024 16:00:50 +0800
| Newsgroups | gmane.comp.ai.weka |
|---|---|
| Message-ID | <CABaQXBsQ7nD1Uvu_EjJQGaBThQHKywfEvrBpKnj735mOtSTaRA@mail.gmail.com> |
--===============5389972718107822315== Content-Type: multipart/alternative; boundary="000000000000f143e506165a5e87" --000000000000f143e506165a5e87 Content-Type: text/plain; charset="UTF-8" Content-Transfer-Encoding: quoted-printable Hello Thank you very much for sharing your information. Our data were obtained from the same PET imaging machine and identical settings. I believe our medical physicists routinely perform QA of high quality for this machine. Previously I was thinking that is the problem of a large number of data set (n>800). So that large number of data set (> 800) actually was not an issue? Shu-Ju Ulrich Mayring <[email protected]> =E6=96=BC 2024=E5=B9=B44=E6=9C=8818= =E6=97=A5 =E9=80=B1=E5=9B=9B =E4=B8=8A=E5=8D=888:59=E5=AF=AB=E9=81=93=EF=BC= =9A > Am 17.04.24 um 05:00 schrieb Shu-Ju Tu: > > Hi dear Weka development team staff: > > > > I have a problem of getting low predictive accuracy when running a larg= e > > data set. > > > > Here is the story and thank you for the patient in advance: > > We started a small data set (n=3D100) last year. > > It is a 2-class supervised data set and the class is evenly distributed > > 50-50. > > The correctly predictive accuracy on training after feature selection > > and test data sets is about 85%. > > We have tried RandomForest and AdaBoostM1. > > Then we increased the data set to n=3D200 (later 300) and were getting > > about similar predictive results. > > Then recently we increased to n=3D800 and were getting very low accurac= y > > of 60%. > > > > Are there something we can do and try to improve on the results? > > Maybe your new data is significantly different from the old data. If so, > you could try to retrain your model on the new data. > > I had a situation like that where I was looking at manufacturing data. > Then they reconfigured / optimised the machine and the data changed > enough to make my model useless. > > > _______________________________________________ > Wekalist mailing list -- [email protected] > Send posts to [email protected] > To unsubscribe send an email to [email protected] > To subscribe, unsubscribe, etc., visit > https://list.waikato.ac.nz/postorius/lists/wekalist.list.waikato.ac.nz > List etiquette: > http://www.cs.waikato.ac.nz/~ml/weka/mailinglist_etiquette.html > --000000000000f143e506165a5e87 Content-Type: text/html; charset="UTF-8" Content-Transfer-Encoding: quoted-printable <div dir=3D"ltr"><div>Hello Thank you very much for sharing your informatio= n.</div><div><br></div><div>Our data were obtained from the same PET imagin= g machine and identical settings.</div><div>I believe our medical physicist= s routinely perform QA of high quality for this machine.</div><div><br></di= v><div>Previously I was thinking that is the problem of a large number of d= ata set (n>800).</div><div>So that large number of data set (> 800) a= ctually was not an issue?</div><div><br></div><div>Shu-Ju<br></div></div><b= r><div class=3D"gmail_quote"><div dir=3D"ltr" class=3D"gmail_attr">Ulrich M= ayring <<a href=3D"mailto:[email protected]">[email protected]= </a>> =E6=96=BC 2024=E5=B9=B44=E6=9C=8818=E6=97=A5 =E9=80=B1=E5=9B=9B = =E4=B8=8A=E5=8D=888:59=E5=AF=AB=E9=81=93=EF=BC=9A<br></div><blockquote clas= s=3D"gmail_quote" style=3D"margin:0px 0px 0px 0.8ex;border-left:1px solid r= gb(204,204,204);padding-left:1ex">Am 17.04.24 um 05:00 schrieb Shu-Ju Tu:<b= r> > Hi dear Weka development team staff:<br> > <br> > I have a problem of getting low predictive accuracy when running a lar= ge <br> > data set.<br> > <br> > Here is the story and thank you for the patient in advance:<br> > We started a small data set (n=3D100) last year.<br> > It is a 2-class supervised data set and the class is evenly distribute= d <br> > 50-50.<br> > The correctly predictive accuracy on training after feature selection = <br> > and test data sets is about 85%.<br> > We have tried RandomForest and AdaBoostM1.<br> > Then we increased the data set to n=3D200 (later 300) and were getting= <br> > about similar predictive results.<br> > Then recently we increased to n=3D800 and were getting very low accura= cy <br> > of 60%.<br> > <br> > Are there something we can do and try to improve on the results?<br> <br> Maybe your new data is significantly different from the old data. If so, <b= r> you could try to retrain your model on the new data.<br> <br> I had a situation like that where I was looking at manufacturing data. <br> Then they reconfigured / optimised the machine and the data changed <br> enough to make my model useless.<br> <br> <br> _______________________________________________<br> Wekalist mailing list -- <a href=3D"mailto:[email protected]" tar= get=3D"_blank">[email protected]</a><br> Send posts to <a href=3D"mailto:[email protected]" target=3D"_bla= nk">[email protected]</a><br> To unsubscribe send an email to <a href=3D"mailto:[email protected]= to.ac.nz" target=3D"_blank">[email protected]</a><br> To subscribe, unsubscribe, etc., visit <a href=3D"https://list.waikato.ac.n= z/postorius/lists/wekalist.list.waikato.ac.nz" rel=3D"noreferrer" target=3D= "_blank">https://list.waikato.ac.nz/postorius/lists/wekalist.list.waikato.a= c.nz</a><br> List etiquette: <a href=3D"http://www.cs.waikato.ac.nz/~ml/weka/mailinglist= _etiquette.html" rel=3D"noreferrer" target=3D"_blank">http://www.cs.waikato= .ac.nz/~ml/weka/mailinglist_etiquette.html</a><br> </blockquote></div> --000000000000f143e506165a5e87-- --===============5389972718107822315== Content-Type: text/plain; charset="us-ascii" MIME-Version: 1.0 Content-Transfer-Encoding: 7bit Content-Disposition: inline _______________________________________________ Wekalist mailing list -- [email protected] Send posts to [email protected] To unsubscribe send an email to [email protected] To subscribe, unsubscribe, etc., visit https://list.waikato.ac.nz/postorius/lists/wekalist.list.waikato.ac.nz List etiquette: http://www.cs.waikato.ac.nz/~ml/weka/mailinglist_etiquette.html --===============5389972718107822315==--