0

我正在使用 PyQuery 处理来自 Web 的大量文档。PyQuery 使用 lxml 来解析 HTML 文档。

事实上,很多文档都不是有效的 HTML。因此,lxml 无法成功解析这些无效文档,这使我无法进一步获取信息。并且经常引发以下异常:

    Traceback (most recent call last):
      File "/usr/local/lib/python2.7/dist-packages/twisted/internet/base.py", line 1201, in mainLoop
        self.runUntilCurrent()
      File "/usr/local/lib/python2.7/dist-packages/twisted/internet/base.py", line 824, in runUntilCurrent
        call.func(*call.args, **call.kw)
      File "/usr/local/lib/python2.7/dist-packages/twisted/internet/defer.py", line 382, in callback
        self._startRunCallbacks(result)
      File "/usr/local/lib/python2.7/dist-packages/twisted/internet/defer.py", line 490, in _startRunCallbacks
        self._runCallbacks()
    --- <exception caught here> ---
      File "/usr/local/lib/python2.7/dist-packages/twisted/internet/defer.py", line 577, in _runCallbacks
        current.result = callback(current.result, *args, **kw)
      File "/home/hxiao/hiit/crawl/crawl/spiders/basic.py", line 40, in parse
        doc = pq(response.body)
      File "/usr/local/lib/python2.7/dist-packages/pyquery/pyquery.py", line 226, in __init__
        elements = fromstring(context, self.parser)
      File "/usr/local/lib/python2.7/dist-packages/pyquery/pyquery.py", line 70, in fromstring
        result = getattr(lxml.html, meth)(context)
      File "/usr/local/lib/python2.7/dist-packages/lxml/html/__init__.py", line 706, in fromstring
        doc = document_fromstring(html, parser=parser, base_url=base_url, **kw)
      File "/usr/local/lib/python2.7/dist-packages/lxml/html/__init__.py", line 600, in document_fromstring
        value = etree.fromstring(html, parser, **kw)
      File "lxml.etree.pyx", line 3032, in lxml.etree.fromstring (src/lxml/lxml.etree.c:68121)

      File "parser.pxi", line 1786, in lxml.etree._parseMemoryDocument (src/lxml/lxml.etree.c:102470)

      File "parser.pxi", line 1674, in lxml.etree._parseDoc (src/lxml/lxml.etree.c:101299)

      File "parser.pxi", line 1074, in lxml.etree._BaseParser._parseDoc (src/lxml/lxml.etree.c:96481)

      File "parser.pxi", line 582, in lxml.etree._ParserContext._handleParseResultDoc (src/lxml/lxml.etree.c:91290)

      File "parser.pxi", line 683, in lxml.etree._handleParseResult (src/lxml/lxml.etree.c:92476)

      File "parser.pxi", line 631, in lxml.etree._raiseParseError (src/lxml/lxml.etree.c:91904)

    lxml.etree.XMLSyntaxError: line 649: htmlParseEntityRef: expecting ';'

我在问什么:

我想要一种让lxml以不那么严格的方式进行解析的方法,以便可以忽略这种无效性。

4

1 回答 1

1

这个答案可能不是很有帮助,但我调查了类似的问题。

也许你可以看看这个 pyquery 的提示?

http://pythonhosted.org/pyquery/tips.html

于 2014-06-04T10:52:41.677 回答