
abstract
- patchTST
- transformer-based
- 다변량 time series forecasting model
- self-supervised representation learning
- component
- transformer에게 token화된 input으로 주기 위해 time series를 subseries level patch로 쪼갬
- 각 channel을 independent single univariate time series로 가정하고, 같은 embedding과 transfomer weight를 모든 series에 대해 공유함.(channel independence)
- 세가지 장점이 있음
- embedding은 local semantic information을 가지고 있음
- attention map의 연산량과 메모리 사용량은 같은 look-back window에서 이차적으로 줄어들음
- model은 더 오래된 과거도 고려할 수 있음
- 따라서 결론
- 다른 transfomer-based model보다 훨씬 좋다.
- large dataset에 대해 supervised training을 했을때, 월등한 fine-tune 성능도 보인다.
- 학습의 전이도 잘됐다
Introduction
forecasting은 time series analysis의 매우 중요한 task중 하나고, 여러 곳에서 좋은 성능을 보이는 deep learning은 forecasting에서도 좋은 성능을 보임. 이건 비단 forecasting뿐만 아니라 classification, 이상탐지와 같은 분야에서도 마찬가지임
그중 transformer는 NLP, CV, speech, time series등 여러 분야에서 great success를 보여왔고, element간의 connection을 자동으로 학습할 수 있는 attention은 squence한 time series에 적합함.
- time series에 transformer 적용한 선행 연구
- informer
- fedformer
- autoformer
- but, 간단한 linear model이 outperform함
- 따라서 patch TST를 통해 개선
- key design
- patching
- time series forecasting은 다른 time step과의 correlation을 이해하는데에 중점이 있음. 그러나 single time step은 언어와 같은 semantic meaning이 없음. 그러므로 local semantic information을 추출하여 서로간의 관계를 파악하는 것이 필수.
- channel-independence
- 다변량을 받아들이는 여러가지 방식이 있음
- channel-mixing: information을 섞기 위해 여러 feature를 하나의 embedding layer로 projection함
- channel-independence: 각 input token은 single channel의 정보밖에 포함하지 않음.
- 기존에는 channel-independence가 CNN과 linear model에서는 잘 작동한다는 것을 확인했으나, trnasformer-based model에 대해서는 아직 적용되지 않았음.
- patching
- patchTST의 장점
- reduction on time and space complexity
- trnaformer의 시간복잡도는 $O(N^2)$이고 여기서 N은 look back window size임.
- 본 모델은 N이 patch의 길이이므로 L / S (L은 look back window, S는 stride) 시간복잡도가 quadratically하게 줄어들음.
- capability of learning from longer look-back window
- look back window를 그냥 늘리는 것은 memory 사용량과 연산량을 너무 증가시킴
- time series는 중복 데이터가 많기 때문에 그걸 줄이기 위해 downsampling, sparse connection of attention과 같은 방법을 사용하고 충분히 잘 forecasting함. 따라서 본 연구에서는 4step씩 건너 뛰며 point를 사용함.(즉, 원본 데이터의 길이가 380일때, input token의 길이는 96)
- 실제로 최근 96개를 이용한거보다 380에서 4step씩 sampling한게 성능이 더 좋음
- 그러면 긴 look-back window를 유지하면서 최대한 value를 throwing하는 것을 피하는 방법은? patching이 좋은 방법
- self-supervised learning 발전
- 선형 모델은 표현력이 너무 낮아서 고차원적인 것을 담을 수 없음capability of representation learning
- reduction on time and space complexity
Related Work
- patch in transformer-based Models
- 모든 application에 대해 patching은 local semantic information이 중요할때 essential한 part임
- NLP: BERT - subword-based tokenization
- CV: Vision Transformer - 이미지를 16*16 patch로 나누어서 transformer model input으로, BEit - patch를 input으로
- speech: sub-sequence level을 input으로
- transformer-based long-term time series forecasting
- LogTrans: convolutional self-attention layers with LogSparse design을 사용. local information을 사용하고 공간 복잡도를 줄이기 위해
- logsparse 참고
- LogTrans: convolutional self-attention layers with LogSparse design을 사용. local information을 사용하고 공간 복잡도를 줄이기 위해

-
- informer: ProbSparse self-attention with distilling technique를 제안. 중요한 key들을 효율적으로 추출하기 위해
- probsparse 참고
- informer: ProbSparse self-attention with distilling technique를 제안. 중요한 key들을 효율적으로 추출하기 위해

-
- autoformer: auto-correlation을 사용. patch level connection을 얻기위해.
- triformer: patch attention 사용. but, 목적은 complexity를 줄이기 위함. 따라서, input unit으로 사용하지도 않고 semantic importance를 발견하지도 않음
- time series represenraion learning
- TST, TS-TCC에서 transformer-based model 사용하지만 아직 잠재력이 실현되지는 않음.
proposed method-model structure
모델 구조

look-back window $L: (x_1, ..., x_L)$에 대해 다변량이므로 각 $x_t$는 $M$개의 dimension을 가지고 있고, $(x_{L+1}, ...x_{L+T})$의 T개의 futer value를 forecast함.
- forward process
- 각 input feature를 1차원으로 나눔. 즉, M개의 look back window
- 그다음 patching
- 그다음 projection + position embedding
- 그다음 trnasformer encoder
- 그다음 flatten + linear head
- 그러면 각 dimention의 futer vector가 추출

- patching
- S: region이 오버랩되지 않도록하는 stride$N=(L-P)/S + 2$, 끝부분을 반복하는 pad S를 추가.
- 연산량, 공간복잡도 quadratically하게 줄일 수 있음.
- patching 결과: $x_p^{(i)}\in\R^{P*N}$
- P: patch length
- transformer encoder
- vanilla transformer encoder 사용.
- 각 patch는 linear projection을 통해 D차원의 latent space로(P → D)
- learnable position encodeing 사용.
- loss functionmse
- MSE 사용.

- instance normalization
- distribution shift effect를 줄이기 위해 사용.
proposed method-representation learning
더 복잡한 representation을 배우기 위해 self-supervise pre-training을 실시. 중간 부분을 masked 하고 그 부분을 복원하도록 하여 학습
masking은 각 time series마다, 다른 채널마다 random하게 됨. 그러나 두가지 잠재적인 issue가 있음
- single time step level에만 masking이 적용됨. 문제는 양옆의 점을 가지고 interpolating을 하여 손쉽게 추론할 수 있음. 따라서 다른 사이즈의 time group을 random하게 masking하여 해결할 수 있음.
- 출력층의 가중치가 너무 많아서 overfitting될 수 있음.
patchTST는 이문제 해결
- prediction head을 D*P linear layer로 바꿈
- non overlapping된 patch를 사용하여 다른 patch가 masked된 정보를 가지고 있지 않도록 함. 그러고 random하게 patch를 고르고 0으로 바꿔서 masked함.
각 채널이 모델을 통과할때 각 변수의 특성에 맞는 latent representation을 만들어냄.
experiments - lon-term time series forecasting
wheather, traffic, electricity, ili, etth1, etth2, ettm1, ettm2 dataset활용
즉, 엄밀히 따지면 Time series foundation model은 아님.
lookback length가 길수록 좋은 성능내는 경향
experiments - representation learning
pre-trained model 에 대해 linear probing과 end-to-end fine-tuning 두 옵션에 대해서 학습하고 평가.
결론
- patch 단위로 입력하면 국소적인 패턴을 잘 이해하고 연산량을 매우 줄일 수 있음.
- 각 변수의 관계를 파악하는 것은 과적합 문제가 있었고, 그냥 독립적인 단변량처럼 처리한 모델이 성능이 더 우세
- 시계열 데이터셋에서도 representation learning이 가능
한계
- channel independence때문에 변수 간의 상관관계를 학습하지 못함.(설명 가능성, OOD 등)
- 연산량을 줄이긴 했어도, transformer이므로 연산량이 많음.(look-back window 길수록)
- foundation이라고 보기는 힘들음.