htdig windows7

John Dorapalli <[email protected]> Wed, 23 Mar 2011 23:22:11 -0400
Newsgroups gmane.comp.web.htdig.general
Message-ID <[email protected]>
--===============2718938943731721851==
Content-Type: multipart/alternative; boundary=00151758f44aa83f4c049f31fb55

--00151758f44aa83f4c049f31fb55
Content-Type: text/plain; charset=ISO-8859-1

Hi,

I am trying to make it work on windows7 for PDF indexing.
All the database files are being generated but I see the following issues,

1) The db.docsdb is generated with pdf id but not with TItle.
2) The excrepts(H) attribute is missing from the db.docs file
3) The db.worddump is generated with junk charecters.

The db.docs and db.worddump files, I tried using the ones generated on linux
which worked fine but not the db.docsdb and db.docs.index files.

Please let me know what options I have?

I tested running perl sccripts doc2html and pdf2html and they are parsing my
pdf but only the local ones. They are not parsing when I pass the URL of the
pdf.
pdftotext and pdfinfo are working fine.

Also, how can index the pdfs in my local system directory.
I tried these options but it didn't work,

start_url:             http://localhost/pdf/
#local_urls:   http://localhost/pdf/ = C:/cygwin/var/www/htdocs/pdf/
#local_urls_only: true


Thanks for your help.

John

--00151758f44aa83f4c049f31fb55
Content-Type: text/html; charset=ISO-8859-1
Content-Transfer-Encoding: quoted-printable

Hi,<br><br>I am trying to make it work on windows7 for PDF indexing.<br>All=
 the database files are being generated but I see the following issues,<br>=
<br>1) The db.docsdb is generated with pdf id but not with TItle.<br>
2) The excrepts(H) attribute is missing from the db.docs file<br>3) The db.=
worddump is generated with junk charecters.<br><br>The
 db.docs and db.worddump files, I tried using the ones generated on=20
linux which worked fine but not the db.docsdb and db.docs.index files.<br>
<br>Please let me know what options I have?<br><br>I tested running perl
 sccripts doc2html and pdf2html and they are parsing my pdf but only the
 local ones. They are not parsing when I pass the URL of the pdf.<br>pdftot=
ext and pdfinfo are working fine.<br>
<br>Also, how can index the pdfs in my local system directory.<br>I tried t=
hese options but it didn&#39;t work,<br><br>start_url:=A0=A0=A0=A0=A0=A0=A0=
=A0=A0=A0=A0=A0 <a href=3D"http://localhost/pdf/" target=3D"_blank">http://=
localhost/pdf/</a><br>#local_urls:=A0=A0 <a href=3D"http://localhost/pdf/" =
target=3D"_blank">http://localhost/pdf/</a> =3D C:/cygwin/var/www/htdocs/pd=
f/<br>

#local_urls_only: true<br><br><br>Thanks for your help.<br><font color=3D"#=
888888"><br>John</font><br>

--00151758f44aa83f4c049f31fb55--


--===============2718938943731721851==
Content-Type: text/plain; charset="us-ascii"
MIME-Version: 1.0
Content-Transfer-Encoding: 7bit
Content-Disposition: inline

------------------------------------------------------------------------------
Enable your software for Intel(R) Active Management Technology to meet the
growing manageability and security demands of your customers. Businesses
are taking advantage of Intel(R) vPro (TM) technology - will your software 
be a part of the solution? Download the Intel(R) Manageability Checker 
today! http://p.sf.net/sfu/intel-dev2devmar
--===============2718938943731721851==
Content-Type: text/plain; charset="us-ascii"
MIME-Version: 1.0
Content-Transfer-Encoding: 7bit
Content-Disposition: inline

_______________________________________________
ht://Dig general mailing list: <[email protected]>
ht://Dig FAQ: http://htdig.sourceforge.net/FAQ.html
List information (subscribe/unsubscribe, etc.)
https://lists.sourceforge.net/lists/listinfo/htdig-general
--===============2718938943731721851==--