我正在尝试评估商店访客对 COVID-19 传播的影响。
这是一个简单的场景:
- 访客 A 走进商店并在时间 = 0 时遇到 Employee1。
- 然后,访客 A 在时间 = 1 处与员工 2 会面。
- 访客 B 走进商店并在时间 = 1 时遇到 Employee1。
- 然后,访客 B 在时间 = 2 处与员工 3 会面。
- 访客 A 离开商店。
当我收集所有访问者数据以及他们在一段时间内遇到的人时,数据集看起来像这样:
表visitorByEmployee
:
| VisitorID | EmployeeID | Contact |
+-----------+------------+-------------------+
| 100 | X123 | 3/11/2020 1:00 |
| 100 | X124 | 3/11/2020 1:10 |
| 101 | X123 | 3/12/2020 1:11 |
| 101 | X125 | 3/11/2020 1:20 |
| 102 | X126 | 3/12/2020 10:00 |
| 102 | X124 | 3/12/2020 10:00 |
| 103 | X123 | 3/12/2020 11:00 |
| 104 | X124 | 3/12/2020 12:00 |
| 104 | X126 | 3/12/2020 12:00 |
| 105 | X126 | 3/12/2020 12:00 |
我想根据这些数据构建一个层次结构,最终可以表示如下:
每棵树都代表了访客对病毒传播的影响:
100
--> X123
--> 101
--> X125
--> 103
--> X124
--> 104
102
--> X126
--> 104
--> 105
--> X124
--> 104
--> X126
我试图通过首先找到根节点(不受之前访问者和/或他们看到的员工影响的根访问者)来做到这一点。这些是 100 和 102。
SELECT
*,
ROW_NUMBER() OVER (PARTITION BY EmployeeID ORDER BY Contact) AS SeenOrder
INTO
#SeenOrder
FROM
visitorByEmployee
SELECT *
INTO #RootVisitors
FROM #SeenOrder
WHERE SeenOrder = 1
从#RootVisitors
和开始#SeenOrder
,我想建立一个表,它可以告诉我影响的层次结构,并可能导致如下结果:
| InitVisitorID | HLevel | EmployeeID | VisitorID |
+---------------+------------+-------------------+-------------+
| 100 | 0 | X123 | 100 |
| 100 | 0 | X124 | 100 |
| 100 | 1 | X123 | 101 |
| 100 | 1 | X123 | 103 |
| 100 | 1 | X124 | 104 |
| 100 | 2 | X125 | 101 |
| 102 | 0 | X126 | 102 |
| 102 | 0 | X124 | 102 |
| 102 | 1 | X126 | 104 |
| 102 | 1 | X126 | 105 |
| 102 | 1 | X124 | 104 |
| 102 | 2 | X126 | 104 |
这是可以使用递归 CTE 完成的吗?我试图这样做,但由于从访客到员工到访客到员工的层次结构不断变化,我很难创建递归 CTE。
更新 这是我正在研究的递归 CTE。它还不起作用,但我正在分享的方法是:
; WITH exposure_tree AS (
/* == Anchor with the root visitors == */
/* == You can think of this: The Employees who were exposed by the Visior == */
SELECT re.VisitorID InitVisitor,
1 as Level,
CASE WHEN 1%2=1 THEN 'Visitor' ELSE 'Employee' END ExposerType,
re.VisitorID Exposer,
re.EmployeeID Exposee,
re.SeenOrder,
re.InitialContact
FROM #SeenOrder re
WHERE re.SeenOrder = 1
/* == Recursive Part #1 ==
Get the visitors who were exposed next by the exposed employees
*/
UNION ALL
SELECT et.VisitorID InitVisitor,
Level + 1,
CASE WHEN (Level+1)%2=1 THEN 'Visitor' ELSE 'Employee' END ExposerType,
re.EmployeeID,
re.VisitorID, -- These are switched from the anchor.
re.SeenOrder,
re.InitialContact
FROM #SeenOrder re
JOIN exposure_tree et ON et.Exposee = re.EmployeeID AND re.SeenOrder > 1 AND re.InitialContact > et.InitialContact
UNION ALL
/* == Recursive Part #2 ==
Get the next employees who were exposed the second level exposed visitors
*/
SELECT et.VisitorID InitVisitor,
Level + 2,
CASE WHEN (Level+2)%2=1 THEN 'Visitor' ELSE 'Employee' END ExposerType,
re.VisitorID,
re.EmployeeID,
re.SeenOrder,
re.InitialContact
FROM #ROOT_EXPOSURES re
JOIN exposure_tree et ON re.VisitorID = et.Exposer and re.SeenOrder > 1 AND re.InitialContact > et.InitialContact
)
select top 1000 * from exposure_tree ORDER BY InitVisitor, Level