Skip to contents

Processes a GBIF occurrence cube (a CSV file or a data frame) into a processed_cube object. Cubes produced by the GBIF cube API can have user-specified column names, so check that your column names match the Darwin Core names expected by this function; if not, supply them with the cols_* arguments. The function stops with an error if it cannot find all required columns.

Usage

process_cube(
  cube_name,
  grid_type = c("automatic", "eea", "mgrs", "eqdgc", "isea3h", "custom", "none"),
  first_year = NULL,
  last_year = NULL,
  force_gridcode = FALSE,
  cols_year = NULL,
  cols_yearMonth = NULL,
  cols_yearMonthDay = NULL,
  cols_cellCode = NULL,
  cols_occurrences = NULL,
  cols_scientificName = NULL,
  cols_minCoordinateUncertaintyInMeters = NULL,
  cols_minTemporalUncertainty = NULL,
  cols_kingdom = NULL,
  cols_family = NULL,
  cols_species = NULL,
  cols_kingdomKey = NULL,
  cols_familyKey = NULL,
  cols_speciesKey = NULL,
  cols_familyCount = NULL,
  cols_sex = NULL,
  cols_lifeStage = NULL,
  separator = NULL
)

Arguments

cube_name

Either the path to a data cube CSV file (e.g. system.file("extdata", "denmark_mammals_cube_eqdgc.csv", package = "b3gbi")) or a data frame containing the cube.

grid_type

(Optional) The grid reference system your cube uses. One of "automatic" (default), "eea", "mgrs", "eqdgc", "isea3h", "custom" or "none". With "automatic" the function attempts to detect the grid from the cell codes and returns an error if it fails. If you want to perform analysis on a cube with custom grid codes (e.g. output from the gcube package) or a cube without grid codes, select "custom" or "none", respectively.

first_year

(Optional) The first year of occurrences to include. If not specified, uses the earliest year present in the cube.

last_year

(Optional) The final year of occurrences to include. If not specified, uses the latest year present in the cube.

force_gridcode

(Optional) Logical. If TRUE, skips the check that cell codes match the expected format of grid_type. Not recommended; invalid codes may cause downstream errors. Default FALSE.

cols_year

(Optional) The name of the column containing the year of occurrence (if something other than 'year'). This column is required unless you have a yearMonth column.

cols_yearMonth

(Optional) The name of the column containing the year and month of occurrence (if present and if other than 'yearMonth'). Use this only if you do not have a year column. The b3gbi package does not use month data, so the function will convert your yearMonth column to a year column.

cols_yearMonthDay

(Optional) The name of the column containing the year, month and day of occurrence (if present and if other than 'yearMonthDay'). Use this only if you do not have year or yearMonth columns. The b3gbi package does not use day or month data, so the function will convert your yearMonthDay column to a year column.

cols_cellCode

(Optional) The name of the column containing the grid reference codes (if other than 'cellCode'). This column is required.

cols_occurrences

(Optional) The name of the column containing the number of occurrences (if other than 'occurrences'). This column is required.

cols_scientificName

(Optional) The name of the column containing the scientific name of the species (if other than 'scientificName'). Note that it is not necessary to have both a species column and a scientificName column. One or the other is sufficient.

cols_minCoordinateUncertaintyInMeters

(Optional) The name of the column containing the minimum coordinate uncertainty of the occurrences (if other than 'minCoordinateUncertaintyInMeters').

cols_minTemporalUncertainty

(Optional) The name of the column containing the minimum temporal uncertainty of the occurrences (if other than 'minTemporalUncertainty').

cols_kingdom

(Optional) The name of the column containing the kingdom the occurring species belongs to (if other than 'kingdom').

cols_family

(Optional) The name of the column containing the family the occurring species belongs to (if other than 'family').

cols_species

(Optional) The name of the column containing the name of the occurring species (if other than 'species'). Note that it is not necessary to have both a species column and a scientificName column. One or the other is sufficient.

cols_kingdomKey

(Optional) The name of the column containing the kingdom key of the occurring species (if other than 'kingdomKey').

cols_familyKey

(Optional) The name of the column containing the family key of the occurring species (if other than 'familyKey').

cols_speciesKey

(Optional) The name of the column containing the species key of the occurring species (if other than 'speciesKey'). The column is required, but note that if you have a 'taxonKey' column you can provide it as the speciesKey.

cols_familyCount

(Optional) The name of the column containing the occurrence count by family.

cols_sex

(Optional) The name of the column containing the sex of the observed individuals.

cols_lifeStage

(Optional) The name of the column containing the life stage of the observed individuals.

separator

(Optional) The column-separating character in your csv file. This should be automatically recognized, so only specify this if you are having trouble.

Value

An object of class processed_cube (or sim_cube when grid_type is "custom" or "none"): a list of metadata (years, number of species, grid type, resolution, ...) plus the processed occurrences in the data element.

Examples

# \donttest{
cube_name <- system.file("extdata", "denmark_mammals_cube_eqdgc.csv",
                         package = "b3gbi")
denmark_example_cube <- process_cube(cube_name)
denmark_example_cube
#> 
#> Processed data cube for calculating biodiversity indicators
#> 
#> Date Range: 1901 - 2024 
#> Single-resolution cube with cell size 0.25degrees 
#> Number of cells: 186 
#> Grid reference system: eqdgc 
#> Coordinate range:
#>  xmin  xmax  ymin  ymax 
#>  3.75 15.25 54.50 58.25 
#> 
#> Total number of observations: 6099 
#> Number of species represented: 54 
#> Number of families represented: 19 
#> 
#> Kingdoms represented: Animalia 
#> 
#> First 10 rows of data (use n = to show more):
#> 
#> # A tibble: 989 × 15
#>     year cellCode  kingdomKey kingdom  familyKey family  taxonKey scientificName
#>    <dbl> <chr>     <chr>      <chr>    <chr>     <chr>   <chr>    <chr>         
#>  1  1901 E010N55CD 1          Animalia 9456      Sciuri… 8211070  Sciurus vulga…
#>  2  1930 E014N55DD 1          Animalia 9456      Sciuri… 8211070  Sciurus vulga…
#>  3  1940 E012N55BA 1          Animalia 9614      Bovidae 2441022  Bos taurus    
#>  4  1943 E011N54BA 1          Animalia 5510      Muridae 2437756  Apodemus flav…
#>  5  1946 E008N56BB 1          Animalia 5510      Muridae 5219833  Micromys minu…
#>  6  1952 E010N57CB 1          Animalia 5510      Muridae 2439261  Rattus norveg…
#>  7  1960 E014N55DD 1          Animalia 5534      Sorici… 8316400  Sorex araneus 
#>  8  1963 E008N57DC 1          Animalia 5534      Sorici… 7571319  Sorex minutus 
#>  9  1968 E014N55DB 1          Animalia 5510      Muridae 2437756  Apodemus flav…
#> 10  1969 E012N55BA 1          Animalia 9701      Canidae 5219243  Vulpes vulpes 
#> # ℹ 979 more rows
#> # ℹ 7 more variables: obs <dbl>, minCoordinateUncertaintyInMeters <dbl>,
#> #   minTemporalUncertainty <dbl>, familyCount <dbl>, xcoord <dbl>,
#> #   ycoord <dbl>, resolution <chr>
# }