R/keep_identify_abundant_per_category.R
keep_abundant_per_category-methods.RdFilters the data to keep only transcripts/genes that are consistently expressed above a threshold across samples. This is a filtering version of identify_abundant_per_category() that removes low-abundance features instead of just marking them.
keep_abundant_per_category(
.data,
abundance = assayNames(.data)[1],
design = NULL,
formula_design = NULL,
minimum_counts = 10,
minimum_proportion = 0.7,
minimum_count_per_million = NULL,
minimum_category = 1,
force = FALSE,
coerce_design = FALSE,
...
)
# S4 method for class 'SummarizedExperiment'
keep_abundant_per_category(
.data,
abundance = assayNames(.data)[1],
design = NULL,
formula_design = NULL,
minimum_counts = 10,
minimum_proportion = 0.7,
minimum_count_per_million = NULL,
minimum_category = 1,
force = FALSE,
coerce_design = FALSE,
...
)
# S4 method for class 'RangedSummarizedExperiment'
keep_abundant_per_category(
.data,
abundance = assayNames(.data)[1],
design = NULL,
formula_design = NULL,
minimum_counts = 10,
minimum_proportion = 0.7,
minimum_count_per_million = NULL,
minimum_category = 1,
force = FALSE,
coerce_design = FALSE,
...
)A `SummarizedExperiment` object containing transcript/gene abundance data
The name of the transcript/gene abundance column (character, preferred)
A design matrix for more complex experimental designs.
A formula to generate the design matrix. Overrides `design` when both are provided
The minimum count threshold for a feature to be considered abundant
The minimum proportion of samples in which a feature must be abundant
...
The minimum number of categories/experimental groups that have sufficient abundance samples.
Should existing .abundant column be replaced; defaults to FALSE.
Should non-categorical design matrix be coerced; defaults to FALSE.
Further arguments taken for compatibility, but are not used.
Returns a filtered version of the input object containing only the features that passed the abundance threshold criteria.
This function provides an alternative filtering solution for edgeR's filterByExpr() function to identify consistently expressed features. A feature is considered abundant if it has CPM > minimum_counts in at least minimum_proportion of samples in at least minimum_category experimental groups (defined by design/formula_design).
This function is similar to identify_abundant_per_category() but in addition to writing to an .abundant column, it also filters out the low-abundance features directly.
Mangiola, S., Molania, R., Dong, R., Doyle, M. A., & Papenfuss, A. T. (2021). tidybulk: an R tidy framework for modular transcriptomic data analysis. Genome Biology, 22(1), 42. doi:10.1186/s13059-020-02233-7
McCarthy, D. J., Chen, Y., & Smyth, G. K. (2012). Differential expression analysis of multifactor RNA-Seq experiments with respect to biological variation. Nucleic Acids Research, 40(10), 4288-4297. DOI: 10.1093/bioinformatics/btp616
## Load airway dataset for examples
data('airway', package = 'airway')
# Ensure a 'condition' column exists for examples expecting it
SummarizedExperiment::colData(airway)$condition <- SummarizedExperiment::colData(airway)$dex
# Basic usage
airway |> keep_abundant_per_category()
#> Warning: Neither formula_design and design provided; build design from formula ~ 1 .
#> class: RangedSummarizedExperiment
#> dim: 15272 8
#> metadata(1): ''
#> assays(1): counts
#> rownames(15272): ENSG00000000003 ENSG00000000419 ... ENSG00000273448
#> ENSG00000273486
#> rowData names(11): gene_id gene_name ... symbol .abundant
#> colnames(8): SRR1039508 SRR1039509 ... SRR1039520 SRR1039521
#> colData names(10): SampleName cell ... BioSample condition
# With custom thresholds
airway |> keep_abundant_per_category(
minimum_counts = 5,
minimum_proportion = 0.5
)
#> Warning: Neither formula_design and design provided; build design from formula ~ 1 .
#> class: RangedSummarizedExperiment
#> dim: 18032 8
#> metadata(1): ''
#> assays(1): counts
#> rownames(18032): ENSG00000000003 ENSG00000000419 ... ENSG00000273487
#> ENSG00000273488
#> rowData names(11): gene_id gene_name ... symbol .abundant
#> colnames(8): SRR1039508 SRR1039509 ... SRR1039520 SRR1039521
#> colData names(10): SampleName cell ... BioSample condition
# Using a factor of interest
airway |> keep_abundant_per_category(formula_design = ~condition)
#> class: RangedSummarizedExperiment
#> dim: 16083 8
#> metadata(1): ''
#> assays(1): counts
#> rownames(16083): ENSG00000000003 ENSG00000000419 ... ENSG00000273486
#> ENSG00000273487
#> rowData names(11): gene_id gene_name ... symbol .abundant
#> colnames(8): SRR1039508 SRR1039509 ... SRR1039520 SRR1039521
#> colData names(10): SampleName cell ... BioSample condition