htdig windows7
John Dorapalli <[email protected]> Wed, 23 Mar 2011 23:22:11 -0400
| Newsgroups | gmane.comp.web.htdig.general |
|---|---|
| Message-ID | <[email protected]> |
--===============2718938943731721851== Content-Type: multipart/alternative; boundary=00151758f44aa83f4c049f31fb55 --00151758f44aa83f4c049f31fb55 Content-Type: text/plain; charset=ISO-8859-1 Hi, I am trying to make it work on windows7 for PDF indexing. All the database files are being generated but I see the following issues, 1) The db.docsdb is generated with pdf id but not with TItle. 2) The excrepts(H) attribute is missing from the db.docs file 3) The db.worddump is generated with junk charecters. The db.docs and db.worddump files, I tried using the ones generated on linux which worked fine but not the db.docsdb and db.docs.index files. Please let me know what options I have? I tested running perl sccripts doc2html and pdf2html and they are parsing my pdf but only the local ones. They are not parsing when I pass the URL of the pdf. pdftotext and pdfinfo are working fine. Also, how can index the pdfs in my local system directory. I tried these options but it didn't work, start_url: http://localhost/pdf/ #local_urls: http://localhost/pdf/ = C:/cygwin/var/www/htdocs/pdf/ #local_urls_only: true Thanks for your help. John --00151758f44aa83f4c049f31fb55 Content-Type: text/html; charset=ISO-8859-1 Content-Transfer-Encoding: quoted-printable Hi,<br><br>I am trying to make it work on windows7 for PDF indexing.<br>All= the database files are being generated but I see the following issues,<br>= <br>1) The db.docsdb is generated with pdf id but not with TItle.<br> 2) The excrepts(H) attribute is missing from the db.docs file<br>3) The db.= worddump is generated with junk charecters.<br><br>The db.docs and db.worddump files, I tried using the ones generated on=20 linux which worked fine but not the db.docsdb and db.docs.index files.<br> <br>Please let me know what options I have?<br><br>I tested running perl sccripts doc2html and pdf2html and they are parsing my pdf but only the local ones. They are not parsing when I pass the URL of the pdf.<br>pdftot= ext and pdfinfo are working fine.<br> <br>Also, how can index the pdfs in my local system directory.<br>I tried t= hese options but it didn't work,<br><br>start_url:=A0=A0=A0=A0=A0=A0=A0= =A0=A0=A0=A0=A0 <a href=3D"http://localhost/pdf/" target=3D"_blank">http://= localhost/pdf/</a><br>#local_urls:=A0=A0 <a href=3D"http://localhost/pdf/" = target=3D"_blank">http://localhost/pdf/</a> =3D C:/cygwin/var/www/htdocs/pd= f/<br> #local_urls_only: true<br><br><br>Thanks for your help.<br><font color=3D"#= 888888"><br>John</font><br> --00151758f44aa83f4c049f31fb55-- --===============2718938943731721851== Content-Type: text/plain; charset="us-ascii" MIME-Version: 1.0 Content-Transfer-Encoding: 7bit Content-Disposition: inline ------------------------------------------------------------------------------ Enable your software for Intel(R) Active Management Technology to meet the growing manageability and security demands of your customers. Businesses are taking advantage of Intel(R) vPro (TM) technology - will your software be a part of the solution? Download the Intel(R) Manageability Checker today! http://p.sf.net/sfu/intel-dev2devmar --===============2718938943731721851== Content-Type: text/plain; charset="us-ascii" MIME-Version: 1.0 Content-Transfer-Encoding: 7bit Content-Disposition: inline _______________________________________________ ht://Dig general mailing list: <[email protected]> ht://Dig FAQ: http://htdig.sourceforge.net/FAQ.html List information (subscribe/unsubscribe, etc.) https://lists.sourceforge.net/lists/listinfo/htdig-general --===============2718938943731721851==--