Problem of large data sets

Shu-Ju Tu <[email protected]> Wed, 17 Apr 2024 11:00:19 +0800
Newsgroups gmane.comp.ai.weka
Message-ID <CABaQXBspPQKLhcvz=00uNgpotZhsNCM0jKFibQ0jdprmGmp+Nw@mail.gmail.com>
--===============6578884727909430551==
Content-Type: multipart/alternative; boundary="000000000000b0bdcc0616420e72"

--000000000000b0bdcc0616420e72
Content-Type: text/plain; charset="UTF-8"

Hi dear Weka development team staff:

I have a problem of getting low predictive accuracy when running a large
data set.

Here is the story and thank you for the patient in advance:
We started a small data set (n=100) last year.
It is a 2-class supervised data set and the class is evenly distributed
50-50.
The correctly predictive accuracy on training after feature selection and
test data sets is about 85%.
We have tried RandomForest and AdaBoostM1.
Then we increased the data set to n=200 (later 300) and were getting about
similar predictive results.
Then recently we increased to n=800 and were getting very low accuracy of
60%.

Are there something we can do and try to improve on the results?

Please advise!

Sincerely yours and warm regards,
Shu-Ju

--000000000000b0bdcc0616420e72
Content-Type: text/html; charset="UTF-8"
Content-Transfer-Encoding: quoted-printable

<div dir=3D"ltr"><div>Hi dear Weka development team staff:</div><div><br></=
div><div>I have a problem of getting low predictive accuracy when running a=
 large data set.</div><div><br></div><div>Here is the story and thank you f=
or the patient in advance:</div><div>We started a small data set (n=3D100) =
last year.</div><div>It is a 2-class supervised data set and the class is e=
venly distributed 50-50.<br></div><div>The correctly predictive accuracy on=
 training after feature selection and test data sets is about 85%.</div><di=
v>We have tried RandomForest and AdaBoostM1.<br></div><div>Then we increase=
d the data set to n=3D200 (later 300) and were getting about similar predic=
tive results.</div><div>Then recently we increased to n=3D800 and=20
were getting very low accuracy of 60%.

</div><div><br></div><div>Are there something we can do and try to improve =
on the results?</div><div><br></div><div>Please advise!</div><div><br></div=
><div>Sincerely yours and warm regards,<br></div><div>Shu-Ju<br></div><div>=
<br></div><div><br></div><div><br></div><div><br></div></div>

--000000000000b0bdcc0616420e72--

--===============6578884727909430551==
Content-Type: text/plain; charset="us-ascii"
MIME-Version: 1.0
Content-Transfer-Encoding: 7bit
Content-Disposition: inline

_______________________________________________
Wekalist mailing list -- [email protected]
Send posts to [email protected]
To unsubscribe send an email to [email protected]
To subscribe, unsubscribe, etc., visit https://list.waikato.ac.nz/postorius/lists/wekalist.list.waikato.ac.nz
List etiquette: http://www.cs.waikato.ac.nz/~ml/weka/mailinglist_etiquette.html

--===============6578884727909430551==--