我想执行大量的配置单元查询并将结果存储在数据框中。
我有一个非常大的数据集,结构如下:
+-------------------+-------------------+---------+--------+--------+
| visid_high| visid_low|visit_num|genderid|count(1)|
+-------------------+-------------------+---------+--------+--------+
|3666627339384069624| 693073552020244687| 24| 2| 14|
|1104606287317036885|3578924774645377283| 2| 2| 8|
|3102893676414472155|4502736478394082631| 1| 2| 11|
| 811298620687176957|4311066360872821354| 17| 2| 6|
|5221837665223655432| 474971729978862555| 38| 2| 4|
+-------------------+-------------------+---------+--------+--------+
我想创建一个派生数据框,它使用每一行作为辅助查询的输入:
result_set = []
for session in sessions.collect()[:100]:
query = "SELECT prop8,count(1) FROM hit_data WHERE dt = {0} AND visid_high = {1} AND visid_low = {2} AND visit_num = {3} group by prop8".format(date,session['visid_high'],session['visid_low'],session['visit_num'])
result = hc.sql(query).collect()
result_set.append(result)
这对一百行按预期工作,但会导致 livy 在更高的负载下超时。
我尝试使用 map 或 foreach:
def f(session):
query = "SELECT prop8,count(1) FROM hit_data WHERE dt = {0} AND visid_high = {1} AND visid_low = {2} AND visit_num = {3} group by prop8".format(date,session.visid_high,session.visid_low,session.visit_num)
return hc.sql(query)
test = sampleRdd.map(f)
导致PicklingError: Could not serialize object: TypeError: 'JavaPackage' object is not callable
. 我从这个答案和这个答案中了解到火花上下文对象不可序列化。
我没有尝试先生成所有查询,然后运行批处理,因为我从这个问题中了解到不支持批处理查询。
我该如何进行?