首页
学习
活动
专区
圈层
工具
发布
社区首页 >问答首页 >solr做网页抓取吗?

问solr做网页抓取吗?
EN

Stack Overflow用户
提问于 2009-11-23 13:24:21
回答 9查看 28.9K关注 0票数 18

我对做网络爬虫很感兴趣。我在看solr。

solr是否做网络爬行,或者做网络爬行的步骤是什么?

EN

回答 9

Stack Overflow用户

发布于 2009-11-23 13:36:00

http://lucene.apache.org/solr/ 5+现在确实可以做网页抓取了!http://lucene.apache.org/solr/

较早的Solr版本不能单独执行web爬行,因为从历史上看,它是一个提供全文搜索功能的搜索服务器。它建立在Lucene之上。

如果您需要使用另一个Solr项目抓取web页面,那么您有许多选择,包括:

http://www.cs.cmu.edu/~rcm/websphinx/

  • JSpider

如果您想使用Lucene或SOLR提供的搜索工具,则需要从web爬行结果构建索引。

另请参阅以下内容:

Lucene crawler (it needs to build lucene index)

票数 20
EN

Stack Overflow用户

发布于 2009-11-23 13:30:13

Solr本身并没有web爬行功能。

是Solr的“实际”爬虫(甚至更多)。

票数 9
EN

Stack Overflow用户

发布于 2016-02-21 00:44:35

Solr5开始支持简单的网络爬行(Java Doc。如果想要搜索,Solr是工具,如果你想抓取,Nutch/Scrapy更好:)

要启动并运行它,您可以详细了解一下here。然而,下面是如何让它在一行中启动和运行:

代码语言:javascript
复制
java 
-classpath <pathtosolr>/dist/solr-core-5.4.1.jar 
-Dauto=yes 
-Dc=gettingstarted     -> collection: gettingstarted
-Ddata=web             -> web crawling and indexing
-Drecursive=3          -> go 3 levels deep
-Ddelay=0              -> for the impatient use 10+ for production
org.apache.solr.util.SimplePostTool   -> SimplePostTool
http://datafireball.com/      -> a testing wordpress blog

这里的爬虫非常“幼稚”,你可以在这里找到this Apache Solr的github repo中的所有代码。

下面是响应的样子:

代码语言:javascript
复制
SimplePostTool version 5.0.0
Posting web pages to Solr url http://localhost:8983/solr/gettingstarted/update/extract
Entering auto mode. Indexing pages with content-types corresponding to file endings xml,json,csv,pdf,doc,docx,ppt,pptx,xls,xlsx,odt,odp,ods,ott,otp,ots,rtf,htm,html,txt,log
SimplePostTool: WARNING: Never crawl an external web site faster than every 10 seconds, your IP will probably be blocked
Entering recursive mode, depth=3, delay=0s
Entering crawl at level 0 (1 links total, 1 new)
POSTed web resource http://datafireball.com (depth: 0)
Entering crawl at level 1 (52 links total, 51 new)
POSTed web resource http://datafireball.com/2015/06 (depth: 1)
...
Entering crawl at level 2 (266 links total, 215 new)
...
POSTed web resource http://datafireball.com/2015/08/18/a-few-functions-about-python-path (depth: 2)
...
Entering crawl at level 3 (846 links total, 656 new)
POSTed web resource http://datafireball.com/2014/09/06/node-js-web-scraping-using-cheerio (depth: 3)
SimplePostTool: WARNING: The URL http://datafireball.com/2014/09/06/r-lattice-trellis-another-framework-for-data-visualization/?share=twitter returned a HTTP result status of 302
423 web pages indexed.
COMMITting Solr index changes to http://localhost:8983/solr/gettingstarted/update/extract...
Time spent: 0:05:55.059

最后,您可以看到所有数据都被正确地编入索引。

票数 5
EN
页面原文内容由Stack Overflow提供。腾讯云小微IT领域专用引擎提供翻译支持
原文链接:

https://stackoverflow.com/questions/1781247

复制
相关文章

相似问题

领券
问题归档专栏文章快讯文章归档关键词归档开发者手册归档开发者手册 Section 归档