继续请教 scrapy 的问题。

class MeizituSpider(scrapy.Spider): name = "meizitu" allowed_domains = ["meizitu.com"] start_urls = ( 'http://www.meizitu.com/', )

def parse(self, response):
    sel = Selector(response)
    for link in sel.xpath('//h2/a/@href').extract():
        request = scrapy.Request(link, callback=self.parse_item)
        yield request

    pages = sel.xpath("//div[@class='navigation']/div[@id='wp_page_numbers']/ul/li/a/@href").extract()
    print('pages: %s' % pages)
    if len(pages) > 2:
        page_link = pages[-2]
        page_link = page_link.replace('/a/', '')
        request = scrapy.Request('http://www.meizitu.com/a/%s' % page_link, callback=self.parse)
        yield request

def parse_item(self, response):
    l = ItemLoader(item=MeizituItem(), response=response)
    l.add_xpath('name', '//h2/a/text()')
    l.add_xpath('tags', "//div[@id='maincontent']/div[@class='postmeta  clearfix']/div[@class='metaRight']/p")
    l.add_xpath('image_urls', "//div[@id='picture']/p/img/@src", Identity())
    l.add_value('url', response.url)
    return l.load_item()
    
    
    
    我找了一份 scrapy 代码进行研究。 
    第一个 parse 抓取的是链接。
    第二个 parse 抓取的是内容。
    我想知道怎么让他把每一个 parse 抓取出来的内容都生成一个输出。我想看一下是什么样子的。以便于自己改造？

laoyur

2016-08-12 11:59:26 +08:00

楼主一人撑起了 Python 节点的半壁江山……

好吧，我也是 Python 新人：）
Scrapy 我也学不久

parse 中，如果 return/yield 出来一个 Request ，那就加入调度器中，等候处理；如果 return/yield 出来一个 item ，那就进 pipelines ，你可以在 pipelines 里面自己对 item 进行处理，进数据库或者啥的。
所以你的问题有点模糊，『输出』到底是什么鬼？自己在 pipelines 中打 log ，或者直接在 parse_item 里面 log ，不都可以吗？

KoleHank

2016-08-12 13:24:48 +08:00

弄一个 jsonpipeline ，然后再 setting.py 里面加上就好了，在 jsonpipeline 里面将抓取的 item 存储起来就可以了

代码类似
```
import json
import codecs

class JsonWriterPipeline(object):

def __init__(self):
self.file = codecs.open('item.json', 'wb','utf-8')

def process_item(self, item, spider):
line = json.dumps(dict(item)) + "\n"
self.file.write(line.decode("unicode_escape"))
return item

```

zanpen2000

2016-08-12 15:59:48 +08:00

我记得是可以自己写中间件进行额外的处理的，要在 settings.py 里的 ITEM_PIPELINES 添加你定义好的处理方法，还要在 spider 的入口处定义你的处理方法，下面的项目是我去年写的，代码比较垃圾，凑合看吧

http://git.oschina.net/zanpen2000/gov_website_crawler_checker

这是一个专为移动设备优化的页面（即为了让你能够在 Google 搜索结果里秒开这个页面），如果你希望参与 V2EX 社区的讨论，你可以继续到 V2EX 上打开本讨论主题的完整版本。

https://www.v2ex.com/t/298843

V2EX 是创意工作者们的社区，是一个分享自己正在做的有趣事物、交流想法，可以遇见新朋友甚至新机会的地方。

V2EX is a community of developers, designers and creative people.