r - 读取多个文件时的内存管理

Question

关于读取多个文件和内存管理有很多问题。我正在寻找可以同时解决这两个问题的信息。

我经常需要将数据的多个部分作为单独的文件读取，将它们重新绑定到一个数据集中，然后对其进行处理。到目前为止，我一直在使用类似下面的东西 - rbinideddataset <- do.call("rbind", lapply(list.files(), read.csv, header = TRUE))

我担心在每种方法中都可以观察到的颠簸。这可能是 rbindeddataset 和 not-yet-rbindeddatasets 同时存在于内存中的实例，但我不知道如何确定。有人可以证实这一点吗？

有什么方法可以将预分配原则扩展到这样的任务？或者其他任何人都知道的技巧可能有助于避免这种颠簸？我还尝试rbindlist了结果lapply，但没有显示凹凸。这是否意味着rbindlist足够聪明来处理这个问题？

data.table 和 Base R 解决方案优于某些软件包的产品。

根据与@Dwin 和@mrip 的讨论，于 2013 年 10 月 7 日编辑

> library(data.table)
> filenames <- list.files()
> 
> #APPROACH 1 #################################
> starttime <- proc.time()
> test <- do.call("rbind", lapply(filenames, read.csv, header = TRUE))
> proc.time() - starttime
   user  system elapsed 
  44.60    1.11   45.98 
> 
> rm(test)
> rm(starttime)
> gc()
          used (Mb) gc trigger   (Mb)  max used   (Mb)
Ncells  350556 18.8     741108   39.6    715234   38.2
Vcells 1943837 14.9  153442940 1170.7 192055310 1465.3
> 
> #APPROACH 2 #################################
> starttime <- proc.time()
> test <- lapply(filenames, read.csv, header = TRUE)
> test2 <- do.call("rbind", test)
> proc.time() - starttime
   user  system elapsed 
  47.09    1.26   50.70 
> 
> rm(test)
> rm(test2)
> rm(starttime)
> gc()
          used (Mb) gc trigger   (Mb)  max used   (Mb)
Ncells  350559 18.8     741108   39.6    715234   38.2
Vcells 1943849 14.9  157022756 1198.0 192055310 1465.3
> 
> 
> #APPROACH 3 #################################
> starttime <- proc.time()
> test <- lapply(filenames, read.csv, header = TRUE)
> test <- do.call("rbind", test)
> proc.time() - starttime
   user  system elapsed 
  48.61    1.93   51.16 
> rm(test)
> rm(starttime)
> gc()
          used (Mb) gc trigger   (Mb)  max used   (Mb)
Ncells  350562 18.8     741108   39.6    715234   38.2
Vcells 1943861 14.9  152965559 1167.1 192055310 1465.3
> 
> 
> #APPROACH 4 #################################
> starttime <- proc.time()
> test <- do.call("rbind", lapply(filenames, fread))

> proc.time() - starttime
   user  system elapsed 
  12.87    0.09   12.95 
> rm(test)
> rm(starttime)
> gc()
          used (Mb) gc trigger  (Mb)  max used   (Mb)
Ncells  351067 18.8     741108  39.6    715234   38.2
Vcells 1964791 15.0  122372447 933.7 192055310 1465.3
> 
> 
> #APPROACH 5 #################################
> starttime <- proc.time()
> test <- do.call("rbind", lapply(filenames, read.csv, header = TRUE))
> proc.time() - starttime
   user  system elapsed 
  51.12    1.62   54.16 
> rm(test)
> rm(starttime)
> gc()
          used (Mb) gc trigger   (Mb)  max used   (Mb)
Ncells  350568 18.8     741108   39.6    715234   38.2
Vcells 1943885 14.9  160270439 1222.8 192055310 1465.3
> 
> 
> #APPROACH 6 #################################
> starttime <- proc.time()
> test <- rbindlist(lapply(filenames, fread ))

> proc.time() - starttime
   user  system elapsed 
  13.62    0.06   14.60 
> rm(test)
> rm(starttime)
> gc()
          used (Mb) gc trigger  (Mb)  max used   (Mb)
Ncells  351078 18.8     741108  39.6    715234   38.2
Vcells 1956397 15.0  128216351 978.3 192055310 1465.3
> 
> 
> #APPROACH 7 #################################
> starttime <- proc.time()
> test <- rbindlist(lapply(filenames, read.csv, header = TRUE))
> proc.time() - starttime
   user  system elapsed 
  48.44    0.83   51.70 
> rm(test)
> rm(starttime)
> gc()
          used (Mb) gc trigger  (Mb)  max used   (Mb)
Ncells  350620 18.8     741108  39.6    715234   38.2
Vcells 1944204 14.9  102573080 782.6 192055310 1465.3

正如预期的那样，fread 节省的时间最多。但是，方法 4,6 和 7 显示最小的内存开销，我不确定为什么。

在此处输入图像描述

score 4 · Accepted Answer

它看起来像rbindlist预先分配内存并一次构建新的数据帧，而do.call(rbind)一次添加一个数据帧，每次都复制它。结果是该rbind方法的运行时间为O(n^2)whilerbindlist以线性时间运行。此外，rbindlist应该避免内存中的颠簸，因为它不必在每次或n迭代期间分配新的数据帧。

一些实验数据：

x<-data.frame(matrix(1:10000,1000,10))
ls<-list()
for(i in 1:10000)
  ls[[i]]<-x+i

rbindtime<-function(i){
  gc()
  system.time(do.call(rbind,ls[1:i]))[3]
}
rbindlisttime<-function(i){
  gc()
  system.time(data.frame(rbindlist(ls[1:i])))[3]
}

ii<-unique(floor(10*1.5^(1:15)))
## [1]   15   22   33   50   75  113  170  256  384  576  864 1297 1946 2919 4378

times<-Vectorize(rbindtime)(ii)
##elapsed elapsed elapsed elapsed elapsed elapsed elapsed elapsed elapsed elapsed 
##  0.009   0.014   0.026   0.049   0.111   0.209   0.350   0.638   1.378   2.645 
##elapsed elapsed elapsed elapsed elapsed 
##  5.956  17.940  30.446  68.033 164.549 

timeslist<-Vectorize(rbindlisttime)(ii)
##elapsed elapsed elapsed elapsed elapsed elapsed elapsed elapsed elapsed elapsed 
##  0.001   0.001   0.001   0.002   0.002   0.003   0.004   0.008   0.009   0.015 
##elapsed elapsed elapsed elapsed elapsed 
##  0.023   0.031   0.046   0.099   0.249

不仅rbindlist速度快得多，尤其是对于长输入，而且运行时间仅线性增加，而do.call(rbind)大约呈二次增长。我们可以通过对每组时间拟合一个对数对数线性模型来确认这一点。

> lm(log(times) ~ log(ii))

Call:
lm(formula = log(times) ~ log(ii))

Coefficients:
(Intercept)      log(ii)  
      -9.73         1.73  

> lm(log(timeslist) ~ log(ii))

Call:
lm(formula = log(timeslist) ~ log(ii))

Coefficients:
(Intercept)      log(ii)  
   -10.0550       0.9455

因此，在实验上，运行时间do.call(rbind)随着 whilen^1.73的增长rbindlist是线性的。

score 2 · Accepted Answer

试试这个：

require(data.table)
system.time({
test3 <- do.call("rbind", lapply(filenames, fread, header = TRUE))
            })

你提到了预分配。fread确实有一个“nrows”参数，但在您事先知道行数的情况下它不会加快其操作（因为它会自动为您预先计算行数，这非常快）。

r - 读取多个文件时的内存管理

2 回答 2

Related

Reference