Problem of large data sets
Shu-Ju Tu <[email protected]> Wed, 17 Apr 2024 11:00:19 +0800
| Newsgroups | gmane.comp.ai.weka |
|---|---|
| Message-ID | <CABaQXBspPQKLhcvz=00uNgpotZhsNCM0jKFibQ0jdprmGmp+Nw@mail.gmail.com> |
--===============6578884727909430551== Content-Type: multipart/alternative; boundary="000000000000b0bdcc0616420e72" --000000000000b0bdcc0616420e72 Content-Type: text/plain; charset="UTF-8" Hi dear Weka development team staff: I have a problem of getting low predictive accuracy when running a large data set. Here is the story and thank you for the patient in advance: We started a small data set (n=100) last year. It is a 2-class supervised data set and the class is evenly distributed 50-50. The correctly predictive accuracy on training after feature selection and test data sets is about 85%. We have tried RandomForest and AdaBoostM1. Then we increased the data set to n=200 (later 300) and were getting about similar predictive results. Then recently we increased to n=800 and were getting very low accuracy of 60%. Are there something we can do and try to improve on the results? Please advise! Sincerely yours and warm regards, Shu-Ju --000000000000b0bdcc0616420e72 Content-Type: text/html; charset="UTF-8" Content-Transfer-Encoding: quoted-printable <div dir=3D"ltr"><div>Hi dear Weka development team staff:</div><div><br></= div><div>I have a problem of getting low predictive accuracy when running a= large data set.</div><div><br></div><div>Here is the story and thank you f= or the patient in advance:</div><div>We started a small data set (n=3D100) = last year.</div><div>It is a 2-class supervised data set and the class is e= venly distributed 50-50.<br></div><div>The correctly predictive accuracy on= training after feature selection and test data sets is about 85%.</div><di= v>We have tried RandomForest and AdaBoostM1.<br></div><div>Then we increase= d the data set to n=3D200 (later 300) and were getting about similar predic= tive results.</div><div>Then recently we increased to n=3D800 and=20 were getting very low accuracy of 60%. </div><div><br></div><div>Are there something we can do and try to improve = on the results?</div><div><br></div><div>Please advise!</div><div><br></div= ><div>Sincerely yours and warm regards,<br></div><div>Shu-Ju<br></div><div>= <br></div><div><br></div><div><br></div><div><br></div></div> --000000000000b0bdcc0616420e72-- --===============6578884727909430551== Content-Type: text/plain; charset="us-ascii" MIME-Version: 1.0 Content-Transfer-Encoding: 7bit Content-Disposition: inline _______________________________________________ Wekalist mailing list -- [email protected] Send posts to [email protected] To unsubscribe send an email to [email protected] To subscribe, unsubscribe, etc., visit https://list.waikato.ac.nz/postorius/lists/wekalist.list.waikato.ac.nz List etiquette: http://www.cs.waikato.ac.nz/~ml/weka/mailinglist_etiquette.html --===============6578884727909430551==--