0

这是重现此问题的代码,但可以通过删除“订单”实体来避免。

import featuretools as ft
import pandas as pd
import numpy as np
df = pd.DataFrame({'member_id': ['AAA', 'AAA',  'AAA', 'AAA', 'AAA',  'JJJ', 'JJJ', 'JJJ'],
                   'order_id': ['0001','0001','0001','0002','0002','1111','1111','1111'],
                   'order_datee': ['2011-01-01','2011-01-01','2011-01-01','2014-01-01','2014-01-01','2013-01-01','2013-01-01','2013-01-01'],
                   'member_join_datee': ['2011-01-01','2011-01-01','2011-01-01','2011-01-01','2011-01-01','2012-01-01','2012-01-01','2012-01-01'],
                   'goods_no':['id1','id2','id3','id4','id5','id6','id7','id8'],
                   'amount': [1,     2,      4,     8,      16,     32,   64,     128],
                   'order_amount': [7,     7,      7,     24,      24,     224,   224,     224],
                   'member_lv': [1,     1,      1,     1,      1,     2,   2,     2]})
df


es = ft.EntitySet(id="abc")
es.entity_from_dataframe("purchases",
                         dataframe = df,
                         index = "purchases_index",
                         time_index = 'order_datee',
                         variable_types = {'order_datee': ft.variable_types.Datetime,
                                           'member_join_datee': ft.variable_types.Datetime,
                                           'amount': ft.variable_types.Numeric,
                                           'order_amount': ft.variable_types.Numeric,
                                           'member_lv': ft.variable_types.Numeric,
                                           })

es.normalize_entity(new_entity_id='members',
                    base_entity_id='purchases',
                    index='member_id',
                    make_time_index = 'member_join_datee',
                    additional_variables=['member_join_datee','member_lv'])

es.normalize_entity(new_entity_id='orders',
                    base_entity_id='purchases',
                    index='order_id',
                    make_time_index = 'order_datee',
                    additional_variables=['order_datee','order_amount'])

fm,features = ft.dfs(entityset=es, target_entity='members')

Traceback (most recent call last):
  File "/.../python3.6/site-packages/featuretools/entityset/entityset.py", line 1204, in _import_from_dataframe
    raise LookupError('Time index not found in dataframe')
4

2 回答 2

2

问题是线路additional_variables=['order_datee','order_amount'])。这会将order_datee列从采购实体移动到订单实体。要将其复制到购买实体而不从订单实体中删除,您应该使用copy_variables. 例如

es.normalize_entity(new_entity_id='orders',
                    base_entity_id='purchases',
                    index='order_id',
                    make_time_index = 'order_datee',
                    copy_variables=["order_datee"],
                    additional_variables=['order_amount'])

在我进行更改后,您的代码将为我运行。

于 2018-12-13T02:25:07.553 回答
-1

从实体“订单”的 time_index 和附加变量中删除“订单日期”后,此问题就消失了。

于 2018-12-13T02:18:21.707 回答