[FFmpeg-devel] lavu/tx: support in-place FFT transforms

This commit adds support for in-place FFT transforms. Since our 
internal transforms were all in-place anyway, this only changes
the permutation on the input.

Unfortunately, research papers were of no help here. All focused
on dry hardware implementations, where permutes are free, or on
software implementations where binary bloat is of no concern so
storing dozen times the transforms for each permutation and version
is not considered bad practice.
Still, for a pure C implementation, it's only around 28% slower
than the multi-megabyte FFTW3 in unaligned mode.

Unlike a closed permutation like with PFA, split-radix FFT bit-reversals
contain multiple NOPs, multiple simple swaps, and a few chained swaps,
so regular single-loop single-state permute loops were not possible.
Instead, we filter out parts of the input indices which are redundant.
This allows for a single branch, and with some clever AVX512 asm,
could possibly be SIMD'd without refactoring.

The inplace_idx array is guaranteed to never be larger than the
revtab array, and in practice only requires around log2(len) entries.

The power-of-two MDCTs can be done in-place as well. And it's
possible to eliminate a copy in the compound MDCTs too, however
it'll be slower than doing them out of place, and we'd need to dirty
the input array.

Patch attached.
Subject: [PATCH] lavu/tx: support in-place FFT transforms

This commit adds support for in-place FFT transforms. Since our
internal transforms were all in-place anyway, this only changes
the permutation on the input.

Unfortunately, research papers were of no help here. All focused
on dry hardware implementations, where permutes are free, or on
software implementations where binary bloat is of no concern so
storing dozen times the transforms for each permutation and version
is not considered bad practice.
Still, for a pure C implementation, it's only around 28% slower
than the multi-megabyte FFTW3 in unaligned mode.

Unlike a closed permutation like with PFA, split-radix FFT bit-reversals
contain multiple NOPs, multiple simple swaps, and a few chained swaps,
so regular single-loop single-state permute loops were not possible.
Instead, we filter out parts of the input indices which are redundant.
This allows for a single branch, and with some clever AVX512 asm,
could possibly be SIMD'd without refactoring.

The inplace_idx array is guaranteed to never be larger than the
revtab array, and in practice only requires around log2(len) entries.

The power-of-two MDCTs can be done in-place as well. And it's
possible to eliminate a copy in the compound MDCTs too, however
it'll be slower than doing them out of place, and we'd need to dirty
the input array.
---
 libavutil/tx.c          | 37 +++++++++++++++++++++++++++++++++++++
 libavutil/tx.h          | 14 +++++++++++++-
 libavutil/tx_priv.h     |  9 ++++++---
 libavutil/tx_template.c | 33 ++++++++++++++++++++++++++++++---
 4 files changed, 86 insertions(+), 7 deletions(-)

Message ID	MTBwHbi--B-2@lynne.ee
State	Accepted
Headers	show Return-Path: <ffmpeg-devel-bounces@ffmpeg.org> X-Original-To: patchwork@ffaux-bg.ffmpeg.org Delivered-To: patchwork@ffaux-bg.ffmpeg.org Received: from ffbox0-bg.mplayerhq.hu (ffbox0-bg.ffmpeg.org [79.124.17.100]) by ffaux.localdomain (Postfix) with ESMTP id A83CE44A637 for <patchwork@ffaux-bg.ffmpeg.org>; Wed, 10 Feb 2021 19:15:58 +0200 (EET) Received: from [127.0.1.1] (localhost [127.0.0.1]) by ffbox0-bg.mplayerhq.hu (Postfix) with ESMTP id 8B2E568A4C5; Wed, 10 Feb 2021 19:15:58 +0200 (EET) X-Original-To: ffmpeg-devel@ffmpeg.org Delivered-To: ffmpeg-devel@ffmpeg.org Received: from w4.tutanota.de (w4.tutanota.de [81.3.6.165]) by ffbox0-bg.mplayerhq.hu (Postfix) with ESMTPS id AF2CB6808A0 for <ffmpeg-devel@ffmpeg.org>; Wed, 10 Feb 2021 19:15:52 +0200 (EET) Received: from w3.tutanota.de (unknown [192.168.1.164]) by w4.tutanota.de (Postfix) with ESMTP id 609171060249 for <ffmpeg-devel@ffmpeg.org>; Wed, 10 Feb 2021 17:15:52 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; q=dns/txt; c=relaxed/relaxed; t=1612977352; s=s1; d=lynne.ee; h=From:From:To:To:Subject:Subject:Content-Description:Content-ID:Content-Type:Content-Type:Content-Transfer-Encoding:Cc:Date:Date:In-Reply-To:MIME-Version:MIME-Version:Message-ID:Message-ID:Reply-To:References:Sender; bh=VdKjBJp85YwZqcpXiThqqwG6m0ER6LYpcc/lTnK5RRk=; b=FLmMXqhNEXVO8LS99NnvFTYOdwTwN1W4PkwCOFMdKDUQDq/HIqdmbfvE/6pkmmtw omDdBtFgZGEcF4ZsfCPTEYdIvmvUy2BOf7l2v4TCOtl5FAhqlqAY4qz9jxel0iAkEPo LdvGADRdEDCIYt7e1wh6xamqAChc8B7sibj64iWwfONNI2UsztIubGbLbt6m2qK1XSb fcGo1Ae2qgiYeeOD8aAN8zj9ppDSpTdWghK956lPrDam2W6hyRoXkF76qmeIRgQWX2Y mZc2IOpa5rj7lpH9sFplldi4GHWy8EklEUGpv8lqB822tW+vUYGKM3EOqfy6AznG3gQ HmW7/7XiVA== Date: Wed, 10 Feb 2021 18:15:52 +0100 (CET) From: Lynne <dev@lynne.ee> To: Ffmpeg Devel <ffmpeg-devel@ffmpeg.org> Message-ID: <MTBwHbi--B-2@lynne.ee> MIME-Version: 1.0 Content-Type: multipart/mixed; boundary="----=_Part_193327_724462963.1612977352200" Subject: [FFmpeg-devel] [PATCH] lavu/tx: support in-place FFT transforms X-BeenThere: ffmpeg-devel@ffmpeg.org X-Mailman-Version: 2.1.20 Precedence: list List-Id: FFmpeg development discussions and patches <ffmpeg-devel.ffmpeg.org> List-Unsubscribe: <https://ffmpeg.org/mailman/options/ffmpeg-devel>, <mailto:ffmpeg-devel-request@ffmpeg.org?subject=unsubscribe> List-Archive: <https://ffmpeg.org/pipermail/ffmpeg-devel> List-Post: <mailto:ffmpeg-devel@ffmpeg.org> List-Help: <mailto:ffmpeg-devel-request@ffmpeg.org?subject=help> List-Subscribe: <https://ffmpeg.org/mailman/listinfo/ffmpeg-devel>, <mailto:ffmpeg-devel-request@ffmpeg.org?subject=subscribe> Reply-To: FFmpeg development discussions and patches <ffmpeg-devel@ffmpeg.org> Errors-To: ffmpeg-devel-bounces@ffmpeg.org Sender: "ffmpeg-devel" <ffmpeg-devel-bounces@ffmpeg.org>
Series	[FFmpeg-devel] lavu/tx: support in-place FFT transforms \| expand [FFmpeg-devel] lavu/tx: support in-place FFT transforms

Context	Check	Description
andriy/x86_make_warn	warning	New warnings during build
andriy/x86_make	success	Make finished
andriy/x86_make_fate	success	Make fate finished
andriy/PPC64_make	success	Make finished
andriy/PPC64_make_fate	success	Make fate finished

[FFmpeg-devel] lavu/tx: support in-place FFT transforms

Checks

Commit Message

Comments

Patch