我对做网络爬虫很感兴趣。我在看solr。
solr是否做网络爬行,或者做网络爬行的步骤是什么?
发布于 2009-11-23 13:36:00
http://lucene.apache.org/solr/ 5+现在确实可以做网页抓取了!http://lucene.apache.org/solr/
较早的Solr版本不能单独执行web爬行,因为从历史上看,它是一个提供全文搜索功能的搜索服务器。它建立在Lucene之上。
如果您需要使用另一个Solr项目抓取web页面,那么您有许多选择,包括:
http://www.cs.cmu.edu/~rcm/websphinx/
如果您想使用Lucene或SOLR提供的搜索工具,则需要从web爬行结果构建索引。
另请参阅以下内容:
发布于 2009-11-23 13:30:13
Solr本身并没有web爬行功能。
是Solr的“实际”爬虫(甚至更多)。
发布于 2016-02-21 00:44:35
Solr5开始支持简单的网络爬行(Java Doc。如果想要搜索,Solr是工具,如果你想抓取,Nutch/Scrapy更好:)
要启动并运行它,您可以详细了解一下here。然而,下面是如何让它在一行中启动和运行:
java
-classpath <pathtosolr>/dist/solr-core-5.4.1.jar
-Dauto=yes
-Dc=gettingstarted -> collection: gettingstarted
-Ddata=web -> web crawling and indexing
-Drecursive=3 -> go 3 levels deep
-Ddelay=0 -> for the impatient use 10+ for production
org.apache.solr.util.SimplePostTool -> SimplePostTool
http://datafireball.com/ -> a testing wordpress blog这里的爬虫非常“幼稚”,你可以在这里找到this Apache Solr的github repo中的所有代码。
下面是响应的样子:
SimplePostTool version 5.0.0
Posting web pages to Solr url http://localhost:8983/solr/gettingstarted/update/extract
Entering auto mode. Indexing pages with content-types corresponding to file endings xml,json,csv,pdf,doc,docx,ppt,pptx,xls,xlsx,odt,odp,ods,ott,otp,ots,rtf,htm,html,txt,log
SimplePostTool: WARNING: Never crawl an external web site faster than every 10 seconds, your IP will probably be blocked
Entering recursive mode, depth=3, delay=0s
Entering crawl at level 0 (1 links total, 1 new)
POSTed web resource http://datafireball.com (depth: 0)
Entering crawl at level 1 (52 links total, 51 new)
POSTed web resource http://datafireball.com/2015/06 (depth: 1)
...
Entering crawl at level 2 (266 links total, 215 new)
...
POSTed web resource http://datafireball.com/2015/08/18/a-few-functions-about-python-path (depth: 2)
...
Entering crawl at level 3 (846 links total, 656 new)
POSTed web resource http://datafireball.com/2014/09/06/node-js-web-scraping-using-cheerio (depth: 3)
SimplePostTool: WARNING: The URL http://datafireball.com/2014/09/06/r-lattice-trellis-another-framework-for-data-visualization/?share=twitter returned a HTTP result status of 302
423 web pages indexed.
COMMITting Solr index changes to http://localhost:8983/solr/gettingstarted/update/extract...
Time spent: 0:05:55.059最后,您可以看到所有数据都被正确地编入索引。

https://stackoverflow.com/questions/1781247
复制相似问题