Aspire specific files with httrack
Solved
Besdu06
-
okuni Posted messages 1325 Status Member -
okuni Posted messages 1325 Status Member -
Hello,
I downloaded a website mirroring software called httrack. However, I would like to only mirror certain pages of the site by specifying, in the options tab, the content (which is always the same for the pages I want to mirror) of the pages. However, I find myself with all the pages of the website.
How can I rectify this? Are there other things to modify in the options tab that I might not have done?
Thank you for all your responses!!!!!^^
Besma
Configuration: Windows 7 / Firefox 4.0.1
I downloaded a website mirroring software called httrack. However, I would like to only mirror certain pages of the site by specifying, in the options tab, the content (which is always the same for the pages I want to mirror) of the pages. However, I find myself with all the pages of the website.
How can I rectify this? Are there other things to modify in the options tab that I might not have done?
Thank you for all your responses!!!!!^^
Besma
Configuration: Windows 7 / Firefox 4.0.1
4 answers
-
I don't know if you can retrieve the filenames you want to download.
I sometimes start from the source code of the pages (displayed with the browser), create the list of files with any editor, and then define it in WinHttrack (from memory, it should be in a box Web Address with the option Download specific files).
Otherwise, as okuni indicates, you can set the maximum depth, in the limit tab.
>>> I have a doubt suddenly, you are talking about Httrack, it's a command line version, why aren't you using WinHttrack?-
-
In my installation (which is quite old), the two programs are installed in the same folder, so they were loaded together.
To be honest, I had never seen that httack.exe existed.
All I can advise you is to download Winhttrack and install it.
It's not complicated to use, its configuration is self-explanatory. -
Oh, I didn't even know there were two different versions. I'm talking about Winhttrack.
The best way to properly mirror is to go through all the options, including the link depth. In your case, you should only set one link depth (or two, because I can't remember if the first one you reference counts as a link depth).
Then, in the filter, you indicate that you only want .doc, .html, or other files, so:
+*.doc +*.html
I think it should work :p
-
-
Hello,
I don't think CCM will give answers since it's a tool for "hacking" a website by retrieving information/content.
https://www.commentcamarche.net/faq/307-devenir-pirate-informatique#q=piratage&cur=2&url=%2F
https://www.commentcamarche.net/infos/25921-pourquoi-ccm-n-aide-pas-la-contrefacon-numerique-des-logiciels/
https://www.commentcamarche.net/infos/25845-charte-d-utilisation-de-commentcamarche-net/
--
If resolved, don't forget to click!-
-
Thank you for the information, Nico, but I took my precautions before doing anything :)
The website scraper is allowed if what is scraped is not used for commercial purposes.... In my case, the scraped site provides documents that can be viewed and downloaded legally, although it must be done document by document, and when there are more than 1500.... it's a bit long to do them one by one.
Thanks again for the info, it's good to know anyway. -
-
Hello everyone,
Having used httrack, I can tell you that this tool is no more intended for hacking than taking a screenshot or using an ftp session. What is forbidden is to reuse them commercially or on a website. Web crawlers were much more useful back in the days of slow dial-up connections. They allow you to retrieve an entire site or a type of documents in just a few clicks. It is even possible to retrieve the images from a site in the image folder of each html page, which does not authorize their usage any more. So let's give credit to Besdu06 for wanting to retrieve documents without having to browse and save hundreds of pages. It is everyone's right to archive and access documents offline. The internet is made for that. There are many other means for a hacker to exercise their guilty talents.
P.S: My use of httrack was too brief for me to respond to the initial request, but I think they are or listed on forums. But I believe it is possible as it seems to me that they have filtering by file type.
Best regards -
-
-
Very nice software :)
When you create a project, you need to go to the options and then filter settings
There you can choose which type of file you take :)
--
Love is like spaghetti; when it's soft, it's cooked. (Belgian proverb)-
-
-
It's simple, to scrape only the pages you want, instead of putting the site address
http://www.example.com
you specify
http://www.example.com/dog
http://www.example.com/cat
http://www.example.com/fish
that way you'll only have the subpages of these pages and not their other pages that talk about food and toys, for example.
so basically, you just specify the scope of your search. -
-
Effectively, I tried to use the mimigenie example but it takes all the pages...
Now the idea of the link depth number is interesting but it's a concept I don't master... How to do it?
I'll give you the link to the site: http://www.legifrance.gouv.fr/affichSarde.do?reprise=true&fastReqId=817860530&idSarde=SARDOBJT000007104919&page=1
I just need to retrieve all the decrees (either html, .doc or .pdf) on electricity production.
Thank you for your patience.^^
-
-
There are options to be configured correctly
it's not the only 'vacuum cleaner' that exists...
and what if you come across a site that is protected???
... ultimately, what is it for???
--
the 'www' is also made for communicating, sharing, and exchanging, right?
thank you for having the courtesy to respond to those who try to help you-
no, it is not a protected site; it is a site where you can freely download documents, but you have to do it one by one, and there are more than 1500 of them. So the objective of httrack is to be able to retrieve these documents, which are either in .doc, .pdf, or HTML format.
Thank you for your help^^
-