about
SGT: A Feature Extraction Function for Sequence Data Mining (arxiv.org)
4 points by dennisy on Mar 23, 2020 | hide | past | pdf | 2 comments on HN

In plain words: A function turns any sequence into a fixed-length list of features that captures both nearby and far-apart patterns, without costing more as the patterns get longer. It grouped and classified sequences more accurately and with less computing than string kernels and LSTM networks.

Abstract · Sequence Graph Transform (SGT): A Feature Embedding Function for Sequence Data Mining

Sequence feature embedding is a challenging task due to the unstructuredness of sequence, i.e., arbitrary strings of arbitrary length. Existing methods are efficient in extracting short-term dependencies but typically suffer from computation issues for the long-term. Sequence Graph Transform (SGT), a feature embedding function, that can extract a varying amount of short- to long-term dependencies without increasing the computation is proposed. SGT's properties are analytically proved for interpretation under normal and uniform distribution assumptions. SGT features yield significantly superior results in sequence clustering and classification with higher accuracy and lower computation as compared to the existing methods, including the state-of-the-art sequence/string Kernels and LSTM.

Chitta Ranjan, Samaneh Ebrahimi, Kamran Paynabar
arXiv:1608.03533 · stat.ML, cs.LG · submitted Aug 11, 2016 · updated Oct 5, 2021
abstract · pdf · html

add comment on HN

didn't see minhash and/or LSH in the list of existing work. As someone not too familiar with the field. Are they not relevant?
Its different information that you can capture in SGT, how all the pairs in your sequence interact, MinHash I believe would help compute the similarity between two sequences based on their union and intersection.