0

我正在尝试使用 Apache Beam 为 Google BigQuery 提供的 I/O API 在本地(Sierra)运行管道。

我按照Beam Python quickstart的建议使用 Virtualenv 设置了我的环境,我可以运行 wordcount.py 示例。我还可以使用beam.Create和正确运行自定义管道beam.ParDo

但我无法使用 BigQuery I/O 运行管道。知道我做错了什么吗?

python脚本如下。

import apache_beam as beam
from apache_beam.utils.pipeline_options import PipelineOptions
from apache_beam.io import WriteToText


class MyDoFn(beam.DoFn):
  def process(self, element):
    return element


def run():
  opts = {
    'project': 'gc-project-name'
  }
  p = beam.Pipeline(options=PipelineOptions(**opts))

  input_query = "SELECT name FROM `gc-project-name.dataset_name.table_name`"

  (p
   | beam.io.Read(beam.io.BigQuerySource(query=input_query))
   | beam.ParDo(MyDoFn())
   | beam.io.WriteToText('output.txt')
  )

  result = p.run()
  result.wait_until_finish()

if __name__ == '__main__':
  run()

当我运行它时,我收到以下错误。

WARNING:root:Task failed: Traceback (most recent call last):
  File "/Users/localuser/Virtualenvs/abeam/lib/python2.7/site-packages/apache_beam/runners/direct/executor.py", line 300, in __call__
result = evaluator.finish_bundle()
  File "/Users/localuser/Virtualenvs/abeam/lib/python2.7/site-packages/apache_beam/runners/direct/transform_evaluator.py", line 208, in finish_bundle
with self._source.reader() as reader:
  File "/Users/localuser/Virtualenvs/abeam/lib/python2.7/site-packages/apache_beam/io/gcp/bigquery.py", line 590, in __enter__
self.client = BigQueryWrapper(client=self.test_bigquery_client)
  File "/Users/localuser/Virtualenvs/abeam/lib/python2.7/site-packages/apache_beam/io/gcp/bigquery.py", line 682, in __init__
self.client = client or bigquery.BigqueryV2(
AttributeError: 'module' object has no attribute 'BigqueryV2'
Traceback (most recent call last):
  File "/Users/localuser/Virtualenvs/abeam/lib/python2.7/site-packages/apache_beam/runners/direct/executor.py", line 300, in __call__
result = evaluator.finish_bundle()
  File "/Users/localuser/Virtualenvs/abeam/lib/python2.7/site-packages/apache_beam/runners/direct/transform_evaluator.py", line 208, in finish_bundle
with self._source.reader() as reader:
  File "/Users/localuser/Virtualenvs/abeam/lib/python2.7/site-packages/apache_beam/io/gcp/bigquery.py", line 590, in __enter__
self.client = BigQueryWrapper(client=self.test_bigquery_client)
  File "/Users/localuser/Virtualenvs/abeam/lib/python2.7/site-packages/apache_beam/io/gcp/bigquery.py", line 682, in __init__
self.client = client or bigquery.BigqueryV2(
AttributeError: 'module' object has no attribute 'BigqueryV2'
WARNING:root:A task failed with exception.
 'module' object has no attribute 'BigqueryV2'
Traceback (most recent call last):
  File "frombigquery.py", line 54, in <module>
run()
  File "frombigquery.py", line 51, in run
result.wait_until_finish()
  File "/Users/localuser/Virtualenvs/abeam/lib/python2.7/site-packages/apache_beam/runners/direct/direct_runner.py", line 157, in wait_until_finish
self._executor.await_completion()
  File "/Users/localuser/Virtualenvs/abeam/lib/python2.7/site-packages/apache_beam/runners/direct/executor.py", line 335, in await_completion
self._executor.await_completion()
  File "/Users/localuser/Virtualenvs/abeam/lib/python2.7/site-packages/apache_beam/runners/direct/executor.py", line 300, in __call__
result = evaluator.finish_bundle()
  File "/Users/localuser/Virtualenvs/abeam/lib/python2.7/site-packages/apache_beam/runners/direct/transform_evaluator.py", line 208, in finish_bundle
with self._source.reader() as reader:
  File "/Users/localuser/Virtualenvs/abeam/lib/python2.7/site-packages/apache_beam/io/gcp/bigquery.py", line 590, in __enter__
self.client = BigQueryWrapper(client=self.test_bigquery_client)
  File "/Users/localuser/Virtualenvs/abeam/lib/python2.7/site-packages/apache_beam/io/gcp/bigquery.py", line 682, in __init__
self.client = client or bigquery.BigqueryV2(
AttributeError: 'module' object has no attribute 'BigqueryV2'
4

1 回答 1

1

安装 Apache Beam Python SDK 时,您必须添加一个附加选项以使用 Google Cloud Platform 相关依赖项。

pip install dist/apache-beam-*.tar.gz[gcp]

于 2017-03-13T23:13:36.163 回答