[PATCH 2/4] media: rockchip: Add JPEG decoder driver
Sascha Hauer <[email protected]>
| Newsgroups | org.infradead.lists.linux-arm-kernel,org.infradead.lists.linux-rockchip,org.kernel.vger.linux-devicetree,org.kernel.vger.linux-kernel,org.kernel.vger.linux-media |
|---|---|
| Message-ID | <[email protected]> |
Add a driver for the JPEG hardware decoder Rockchip integrates into a number of its SoCs. Downstream it is known as the VDPU720, which is the name the register definitions carry here. It is a Rockchip in-house core rather than a Verisilicon one, so it gets its own driver next to the other Rockchip media blocks instead of joining the hantro variants. Exposes one V4L2 M2M device implementing the stateful decoder interface: JPEG input -> NV12 output. The frame header is parsed on the CPU when a coded buffer is queued, which is where the resolution for V4L2_EVENT_SOURCE_CHANGE comes from; userspace cannot know the frame size without parsing the bitstream itself, so it learns it from the event and sizes its capture buffers accordingly. V4L2_DEC_CMD_STOP drains and raises V4L2_EVENT_EOS. Hardware requirements: - a DMA side buffer with Q-tables (zigzag->raster), Huffman mincode and value tables, rebuilt from the frame header on every run - a two-phase IRQ clear sequence, which some revisions require. It is done unconditionally, see rkjpegd-vdpu720-regs.h - 16-byte stream alignment with a start-byte offset - MCU-aligned PIC_H (e.g. 1080p YUV420 needs 1088, not 1080) - FILL_DOWN_E on every NV12 conversion, the output chroma is vertically subsampled so the hardware completes the bottom of the picture - DRI (restart interval) support when present - a neutral chroma plane written by the driver for a grayscale frame. The output format converter has no YUV400 path, so the hardware writes the luma plane and leaves the chroma alone, which comes out green Every address register is a single 32 bit one with no high order companion anywhere in the map, so the DMA masks are capped accordingly. That is no constraint in practice: the block sits behind the Rockchip IOMMU, whose address space is 32 bit wide by construction, and memory above 4GB is reached through it. A frame that carries no EOI marker is rejected: that is what a coded buffer too small for the frame looks like, and the hardware would otherwise hand out a half decoded picture indistinguishable from a good one. A frame larger than the negotiated capture format is rejected too, since its dimensions would be programmed against strides taken from that format. Error recovery tries the in-block soft reset first and falls back to pulsing the reset lines. The interrupt handler decodes the error status registers into a readable diagnostic. Assisted-by: Claude:claude-opus-5 --- MAINTAINERS | 1 + drivers/media/platform/rockchip/Kconfig | 1 + drivers/media/platform/rockchip/Makefile | 1 + drivers/media/platform/rockchip/rkjpegd/Kconfig | 14 + drivers/media/platform/rockchip/rkjpegd/Makefile | 3 + .../rockchip/rkjpegd/rkjpegd-vdpu720-regs.h | 230 ++ drivers/media/platform/rockchip/rkjpegd/rkjpegd.c | 2279 ++++++++++++++++++++ 7 files changed, 2529 insertions(+) diff --git a/MAINTAINERS b/MAINTAINERS index baea7e42663a7..9bc3dda5a4b09 100644 --- a/MAINTAINERS +++ b/MAINTAINERS @@ -23431,6 +23431,7 @@ L: [email protected] L: [email protected] S: Maintained F: Documentation/devicetree/bindings/media/rockchip,jpeg-decoder.yaml +F: drivers/media/platform/rockchip/rkjpegd/ ROCKCHIP RK3568 RANDOM NUMBER GENERATOR SUPPORT M: Daniel Golle <[email protected]> diff --git a/drivers/media/platform/rockchip/Kconfig b/drivers/media/platform/rockchip/Kconfig index ba401d32f01ba..53d28eb89e472 100644 --- a/drivers/media/platform/rockchip/Kconfig +++ b/drivers/media/platform/rockchip/Kconfig @@ -5,4 +5,5 @@ comment "Rockchip media platform drivers" source "drivers/media/platform/rockchip/rga/Kconfig" source "drivers/media/platform/rockchip/rkcif/Kconfig" source "drivers/media/platform/rockchip/rkisp1/Kconfig" +source "drivers/media/platform/rockchip/rkjpegd/Kconfig" source "drivers/media/platform/rockchip/rkvdec/Kconfig" diff --git a/drivers/media/platform/rockchip/Makefile b/drivers/media/platform/rockchip/Makefile index 0e0b2cbbd4bd2..c256c51ced269 100644 --- a/drivers/media/platform/rockchip/Makefile +++ b/drivers/media/platform/rockchip/Makefile @@ -2,4 +2,5 @@ obj-y += rga/ obj-y += rkcif/ obj-y += rkisp1/ +obj-y += rkjpegd/ obj-y += rkvdec/ diff --git a/drivers/media/platform/rockchip/rkjpegd/Kconfig b/drivers/media/platform/rockchip/rkjpegd/Kconfig new file mode 100644 index 0000000000000..4e2b40ef60d83 --- /dev/null +++ b/drivers/media/platform/rockchip/rkjpegd/Kconfig @@ -0,0 +1,14 @@ +# SPDX-License-Identifier: GPL-2.0 +config VIDEO_ROCKCHIP_JPEGD + tristate "Rockchip JPEG decoder driver" + depends on ARCH_ROCKCHIP || COMPILE_TEST + depends on VIDEO_DEV + select MEDIA_CONTROLLER + select V4L2_JPEG_HELPER + select V4L2_MEM2MEM_DEV + select VIDEOBUF2_DMA_CONTIG + help + Support for the JPEG decoder Rockchip integrates into a number of + its SoCs, decoding JPEG and MJPEG frames to NV12. + To compile this driver as a module, choose M here: the module + will be called rockchip-jpegd. diff --git a/drivers/media/platform/rockchip/rkjpegd/Makefile b/drivers/media/platform/rockchip/rkjpegd/Makefile new file mode 100644 index 0000000000000..2aaf6bfcfdb74 --- /dev/null +++ b/drivers/media/platform/rockchip/rkjpegd/Makefile @@ -0,0 +1,3 @@ +obj-$(CONFIG_VIDEO_ROCKCHIP_JPEGD) += rockchip-jpegd.o + +rockchip-jpegd-y += rkjpegd.o diff --git a/drivers/media/platform/rockchip/rkjpegd/rkjpegd-vdpu720-regs.h b/drivers/media/platform/rockchip/rkjpegd/rkjpegd-vdpu720-regs.h new file mode 100644 index 0000000000000..5e2cb017cca51 --- /dev/null +++ b/drivers/media/platform/rockchip/rkjpegd/rkjpegd-vdpu720-regs.h @@ -0,0 +1,230 @@ +/* SPDX-License-Identifier: GPL-2.0 */ +/* + * Rockchip VPU720 JPEG decoder register definitions + * + * Derived from downstream Rockchip MPP HAL (hal_jpegd_rkv_reg.h). + * Copyright (C) 2020 Rockchip Electronics Co., Ltd. + * Copyright (C) 2026 WolfVision GmbH + */ +#ifndef RKJPEGD_VDPU720_REGS_H_ +#define RKJPEGD_VDPU720_REGS_H_ + +#include <linux/bitfield.h> +#include <linux/bits.h> +#include <linux/align.h> +#include <linux/types.h> + +#define VDPU720_REG_VERSION 0x000 +#define VDPU720_PROD_NUM GENMASK(31, 16) +#define VDPU720_BIT_DEPTH BIT(8) + +#define VDPU720_REG_INT 0x004 +#define VDPU720_DEC_E BIT(0) +#define VDPU720_IRQ_DIS BIT(1) +#define VDPU720_TIMEOUT_E BIT(2) +#define VDPU720_BUF_EMPTY_E BIT(3) +#define VDPU720_BUF_EMPTY_RELOAD BIT(4) +#define VDPU720_SOFT_RST_EN BIT(5) +#define VDPU720_IRQ_RAW BIT(6) +#define VDPU720_WAIT_RESET_E BIT(7) +#define VDPU720_IRQ BIT(8) +#define VDPU720_DEC_RDY BIT(9) +#define VDPU720_BUS_ERR BIT(10) +#define VDPU720_DEC_ERR BIT(11) +#define VDPU720_TIMEOUT BIT(12) +#define VDPU720_BUF_EMPTY BIT(13) +#define VDPU720_SOFT_RST_RDY BIT(14) + +/* Error status bits used to decide whether a hardware reset is needed */ +#define VDPU720_ERR_MASK (VDPU720_BUS_ERR | VDPU720_DEC_ERR | \ + VDPU720_TIMEOUT | VDPU720_BUF_EMPTY) + +/* + * VPU720 IRQ clear mask. + * + * Some revisions require a two-step IRQ acknowledgment before checking + * VDPU720_IRQ_RAW. Compute the masked value and write it back, then + * write 0 to fully clear. The downstream BSP uses: + * + * clr = (~(0x00fe7f40 & status)) & (0xff0180bf & status) + * + * Preserve "status" bits from 0xff0180bf while clearing only those + * bits NOT already set in 0x00fe7f40. + * + * It is applied unconditionally here. The BSP gates it on the revision in + * VDPU720_REG_VERSION matching 0xdb1f0006 exactly, and the RK3588 reports + * 0xdb1f0005, so it is not needed on that part. The extra write is + * harmless there and saves a revision check that would need silicon we do + * not have to validate. + */ +#define VDPU720_IRQ_CLR_KEEP 0xff0180bf +#define VDPU720_IRQ_CLR_COND 0x00fe7f40 + +#define VDPU720_REG_SYS 0x008 +#define VDPU720_FORCE_SOFTRST BIT(17) /* set when dec_e=0 before soft-reset */ +#define VDPU720_FILL_DOWN_E BIT(24) /* fill bottom padding rows */ +#define VDPU720_FILL_RIGHT_E BIT(25) +#define VDPU720_OUT_SEQ BIT(26) /* 0=raster, 1=tile */ +#define VDPU720_YUV_OUT_FMT GENMASK(29, 27) +#define VDPU720_YUV_OUT_FMT_NATIVE 0 /* no format conversion */ +#define VDPU720_YUV_OUT_FMT_NV12 3 /* output as NV12 */ + +#define VDPU720_REG_PIC_SIZE 0x00c +#define VDPU720_PIC_W_M1 GENMASK(15, 0) +#define VDPU720_PIC_H_M1 GENMASK(31, 16) + +#define VDPU720_REG_PIC_FMT 0x010 +#define VDPU720_JPEG_MODE GENMASK(2, 0) +#define VDPU720_JPEG_MODE_YUV400 0 +#define VDPU720_JPEG_MODE_YUV411 1 +#define VDPU720_JPEG_MODE_YUV420 2 +#define VDPU720_JPEG_MODE_YUV422 3 +#define VDPU720_JPEG_MODE_YUV440 4 +#define VDPU720_JPEG_MODE_YUV444 5 +#define VDPU720_PIX_DEPTH GENMASK(6, 4) +#define VDPU720_PIX_DEPTH_8 0 +#define VDPU720_PIX_DEPTH_12 1 +/* qtables_sel: number of Q-table sets (0..3) */ +#define VDPU720_QTBL_SEL GENMASK(9, 8) +/* htables_sel: number of H-table sets (0..3) */ +#define VDPU720_HTBL_SEL GENMASK(13, 12) +/* dri_e: restart interval enable */ +#define VDPU720_DRI_E BIT(15) +/* dri_mcu_num_m1: restart interval MCU count minus 1 */ +#define VDPU720_DRI_MCU_M1 GENMASK(31, 16) + +#define VDPU720_REG_HOR_STRIDE 0x014 +#define VDPU720_Y_HOR_STRIDE GENMASK(15, 0) +#define VDPU720_UV_HOR_STRIDE GENMASK(31, 16) + +#define VDPU720_REG_Y_VSTRIDE 0x018 +#define VDPU720_Y_VSTRIDE GENMASK(31, 4) + +#define VDPU720_REG_TBL_LEN 0x01c +#define VDPU720_QTBL_LEN GENMASK(4, 0) +#define VDPU720_HTBL_MINCODE_LEN GENMASK(12, 8) +#define VDPU720_HTBL_VALUE_LEN GENMASK(21, 16) +/* bit 16 of the Y horizontal stride, low 16 bits live in REG5 */ +#define VDPU720_Y_HOR_STRIDE_H BIT(24) + +#define VDPU720_REG_STRM_LEN 0x020 +#define VDPU720_STRM_START_BYTE GENMASK(3, 0) +#define VDPU720_STRM_LEN GENMASK(31, 4) + +#define VDPU720_REG_QTBL_BASE 0x024 /* Q-table side buffer, 64-byte aligned */ +#define VDPU720_REG_HTBL_MINCODE 0x028 /* H-mincode table, 64-byte aligned */ +#define VDPU720_REG_HTBL_VALUE 0x02c /* H-value table, 64-byte aligned */ +#define VDPU720_REG_STRM_BASE 0x030 /* JPEG entropy stream, 16-byte aligned */ +#define VDPU720_REG_OUT_BASE 0x034 /* NV12 output buffer, 64-byte aligned */ + +#define VDPU720_REG_STRM_ERR 0x038 +#define VDPU720_ERROR_PRC_MODE BIT(0) +#define VDPU720_STRM_FFFF_ERR_MODE GENMASK(6, 5) +#define VDPU720_STRM_OTHER_MODE GENMASK(8, 7) +/* Recommended default: accept errors, skip 0xFFFF, skip unknown markers */ +#define VDPU720_STRM_ERR_DFLT (VDPU720_ERROR_PRC_MODE | \ + FIELD_PREP_CONST(VDPU720_STRM_FFFF_ERR_MODE, 2) | \ + FIELD_PREP_CONST(VDPU720_STRM_OTHER_MODE, 2)) + +#define VDPU720_REG_CLK_GATE 0x040 +#define VDPU720_CLK_GATE_ALL 0xff + +#define VDPU720_REG_PERF_CTRL 0x078 +#define VDPU720_PERF_WORK_E BIT(0) +#define VDPU720_PERF_CLR_E BIT(1) +#define VDPU720_PERF_CNT_TYPE BIT(3) +#define VDPU720_PERF_RD_LAT_ID GENMASK(7, 4) + +#define VDPU720_REG_AXI_CFG 0x07c +#define VDPU720_ADDR_ALIGN_TYPE GENMASK(1, 0) +#define VDPU720_AR_CNT_ID_TYPE BIT(2) /* 1 = count sw_ar_count_id only */ +#define VDPU720_AW_CNT_ID_TYPE BIT(3) /* 1 = count sw_aw_count_id only */ +#define VDPU720_AR_COUNT_ID GENMASK(7, 4) +#define VDPU720_AW_COUNT_ID GENMASK(11, 8) +#define VDPU720_RD_TOTAL_BYTES_MODE BIT(12) /* 1 = count sw_ar_count_id bytes only */ + +#define VDPU720_REG_DBG_MCU_POS 0x080 +#define VDPU720_DBG_MCU_POS_X GENMASK(15, 0) /* column in MCU units */ +#define VDPU720_DBG_MCU_POS_Y GENMASK(31, 16) /* row in MCU units */ + +#define VDPU720_REG_DBG_ERROR 0x084 +#define VDPU720_DERR_DRI_SEQ BIT(0) /* DRI not at expected sequence */ +#define VDPU720_DERR_STREAM_R0 BIT(1) /* special marker 0 detected */ +#define VDPU720_DERR_STREAM_R1 BIT(2) /* special marker 1 detected */ +#define VDPU720_DERR_STREAM_FFFF BIT(3) /* 0xFFFF sequence in stream */ +#define VDPU720_DERR_OTHER_MARK BIT(4) /* unknown JPEG marker */ +#define VDPU720_DERR_MCU_CNT_L BIT(8) /* restart mark arrived too early */ +#define VDPU720_DERR_MCU_CNT_M BIT(9) /* restart mark arrived too late */ +#define VDPU720_DERR_EOI_NO_END BIT(10) /* EOI before frame complete */ +#define VDPU720_DERR_END_NO_EOI BIT(11) /* frame complete without EOI */ +#define VDPU720_DERR_OVERFLOW BIT(12) /* Huffman coefficient overflow */ +#define VDPU720_DERR_HUFF_EMPTY BIT(13) /* bitstream empty before EOI */ +#define VDPU720_DERR_FLAGS GENMASK(13, 0) /* all of the above */ +#define VDPU720_DERR_FIRST_IDX GENMASK(19, 16) /* index of first error */ + +#define VDPU720_REG_PERF_RD_MAX_LAT 0x088 /* peak read latency (clock cycles) */ +#define VDPU720_REG_PERF_RD_LAT_SAMP 0x08c /* read transactions above lat threshold */ +#define VDPU720_REG_PERF_RD_LAT_ACC 0x090 /* accumulated read latency sum */ +#define VDPU720_REG_PERF_RD_BYTES 0x094 /* total AXI read bytes this frame */ +#define VDPU720_REG_PERF_WR_BYTES 0x098 /* total AXI write bytes this frame */ + +#define VDPU720_REG_PERF_CYCLES 0x09c + +/* + * Side-buffer layout for Q-tables and Huffman tables + * + * The VPU720 JPEG decoder reads quantisation tables and Huffman + * tables from a contiguous DMA buffer with the following layout: + * + * [0, QTBL_SIZE): Q-table data (u16, raster-scan order) + * [HMINCODE_OFF, +HMIN_SZ): Huffman mincode table + * [HVALUE_OFF, +HVAL_SZ): Huffman value table + * + * The Q-tables are per component, one entry each + */ +#define VDPU720_NB_COMPONENTS 3 +#define VDPU720_QTBL_ENTRIES 64 /* 64 coefficients per table */ +#define VDPU720_QTBL_COMP_SIZE (VDPU720_QTBL_ENTRIES * sizeof(u16)) +#define VDPU720_QTBL_SIZE (VDPU720_QTBL_COMP_SIZE * VDPU720_NB_COMPONENTS) + +/* + * The Huffman tables are not per component. The hardware holds two sets and + * has no per component selector register, so the mapping is fixed: the first + * set is used for the luma component and the second one for both chroma + * components. A grayscale frame only needs the first. + */ +#define VDPU720_NB_HTBL_SETS 2 + +/* + * Per-set mincode layout: 16 DC mincodes + 8 DC accaddr pairs + + * 16 AC mincodes + 8 AC accaddr pairs = 48 u16 = 96 bytes + */ +#define VDPU720_HMINCODE_SET_SIZE (48 * sizeof(u16)) +#define VDPU720_HMINCODE_SIZE (VDPU720_HMINCODE_SET_SIZE * VDPU720_NB_HTBL_SETS) +#define VDPU720_HMINCODE_OFF VDPU720_QTBL_SIZE + +/* Per-set value layout: 16 DC values + 176 AC values = 192 bytes */ +#define VDPU720_HVALUE_SET_SIZE 192 +#define VDPU720_HVALUE_SIZE (VDPU720_HVALUE_SET_SIZE * VDPU720_NB_HTBL_SETS) +#define VDPU720_HVALUE_OFF (VDPU720_HMINCODE_OFF + \ + ALIGN(VDPU720_HMINCODE_SIZE, 64)) + +#define VDPU720_TABLE_BUF_SIZE (VDPU720_HVALUE_OFF + VDPU720_HVALUE_SIZE) + +/* + * The three table length registers count 16 byte units, minus one. Derive + * them from the sizes above so that what is programmed always matches what + * the driver writes into the side buffer. + */ +#define VDPU720_TBL_LEN_UNIT 16 +#define VDPU720_TBL_LEN(bytes) ((bytes) / VDPU720_TBL_LEN_UNIT - 1) + +/* + * Huffman value sub-layout per set (192 bytes total): + * bytes [0..15]: DC code values (up to 12 valid entries) + * bytes [16..191]: AC code values (up to 162 valid entries) + */ +#define VDPU720_DC_VALUES_MAX 16 +#define VDPU720_AC_VALUES_MAX 176 /* 12*16 - 16 */ + +#endif /* RKJPEGD_VDPU720_REGS_H_ */ diff --git a/drivers/media/platform/rockchip/rkjpegd/rkjpegd.c b/drivers/media/platform/rockchip/rkjpegd/rkjpegd.c new file mode 100644 index 0000000000000..cccf365ebc5d4 --- /dev/null +++ b/drivers/media/platform/rockchip/rkjpegd/rkjpegd.c @@ -0,0 +1,2279 @@ +// SPDX-License-Identifier: GPL-2.0 +/* + * Rockchip JPEG decoder driver + * + * The register programming is ported from the Rockchip MPP HAL + * (hal_jpegd_rkv.c / hal_jpegd_vpu7xx_com.c) and the downstream + * mpp_jpgdec.c kernel driver. + * + * Copyright (C) 2020 Rockchip Electronics Co., Ltd. + * Copyright (C) 2026 WolfVision GmbH + * Author: Lucas Sinn <[email protected]> + * Copyright (C) 2026 Pengutronix, Sascha Hauer <[email protected]> + */ + +#include <linux/align.h> +#include <linux/bitfield.h> +#include <linux/clk.h> +#include <linux/delay.h> +#include <linux/dma-mapping.h> +#include <linux/interrupt.h> +#include <linux/io.h> +#include <linux/iopoll.h> +#include <linux/log2.h> +#include <linux/module.h> +#include <linux/platform_device.h> +#include <linux/pm_runtime.h> +#include <linux/reset.h> +#include <linux/slab.h> +#include <linux/videodev2.h> +#include <linux/workqueue.h> + +#include <media/media-device.h> +#include <media/v4l2-device.h> +#include <media/v4l2-event.h> +#include <media/v4l2-fh.h> +#include <media/v4l2-ioctl.h> +#include <media/v4l2-jpeg.h> +#include <media/v4l2-mem2mem.h> +#include <media/videobuf2-core.h> +#include <media/videobuf2-dma-contig.h> + +#include "rkjpegd-vdpu720-regs.h" + +#define RKJPEGD_NAME "rockchip-jpegd" + +/* + * The reference manual gives a minimum of 48x48 and a maximum of 65536x65536 + * with a step of 8 pixels, but the upper end of that range is not usable: a + * JPEG frame header cannot describe more than 65535 pixels, the Y_VSTRIDE + * register field runs out at roughly 46340 square, and sizeimage is computed + * as width * height * 2 in 32 bits, which wraps beyond the same point. Cap + * the advertised size well below all three, still four times 4K in each + * direction. + */ +#define RKJPEGD_MIN_WIDTH 48 +#define RKJPEGD_MIN_HEIGHT 48 +#define RKJPEGD_MAX_SIZE 16384 + +/* The coded queue steps in MCU-sized units, the raw one in macroblocks. */ +#define RKJPEGD_CODED_STEP 8 +#define RKJPEGD_RAW_STEP 16 + +/* Bytes per pixel used to size a coded buffer for a given resolution. */ +#define RKJPEGD_CODED_MAX_DEPTH 2 + +/* Milliseconds a frame may take before the watchdog resets the block. */ +#define RKJPEGD_TIMEOUT_MS 2000 + +static const char * const rkjpegd_clk_names[] = { + "aclk", "hclk", +}; + +#define RKJPEGD_NUM_CLOCKS ARRAY_SIZE(rkjpegd_clk_names) + +/** + * struct rkjpegd_aux_buf - auxiliary DMA buffer for hardware tables + * + * @cpu: CPU pointer to the buffer. + * @dma: DMA address of the buffer. + * @size: Size of the buffer in bytes. + */ +struct rkjpegd_aux_buf { + void *cpu; + dma_addr_t dma; + size_t size; +}; + +/** + * struct rkjpegd_src_buf - a coded buffer and the header parsed out of it + * + * @base: videobuf2 mem2mem buffer, must be first. + * @header: header as returned by v4l2_jpeg_parse_header(). + * @scan: storage @header.scan points at. + * @quantization_tables: storage @header.quantization_tables points at. + * @huffman_tables: storage @header.huffman_tables points at. + * @parsed: @header describes a frame this hardware can decode. + * + * The references in @header point into the buffer payload, which stays + * mapped for as long as the buffer is queued, so the header parsed when the + * buffer arrives is still usable when the job runs. + */ +struct rkjpegd_src_buf { + struct v4l2_m2m_buffer base; + struct v4l2_jpeg_header header; + struct v4l2_jpeg_scan_header scan; + struct v4l2_jpeg_reference quantization_tables[4]; + struct v4l2_jpeg_reference huffman_tables[4]; + bool parsed; +}; + +static inline struct rkjpegd_src_buf * +vb2_to_rkjpegd_src_buf(struct vb2_buffer *vb) +{ + return container_of(to_vb2_v4l2_buffer(vb), struct rkjpegd_src_buf, + base.vb); +} + +/** + * struct rkjpegd_dev - the decoder device + * + * @v4l2_dev: V4L2 device. + * @mdev: media device. + * @vdev: video device. + * @m2m_dev: mem2mem device. + * @dev: driver model device. + * @clocks: clocks named by @rkjpegd_clk_names. + * @resets: the block's reset lines, as one array control. + * @regs: register window. + * @vdev_lock: serialises ioctls and the videobuf2 queues. + * @watchdog_work: fires when a job does not complete in time. + */ +struct rkjpegd_dev { + struct v4l2_device v4l2_dev; + struct media_device mdev; + struct video_device vdev; + struct v4l2_m2m_dev *m2m_dev; + struct device *dev; + struct clk_bulk_data clocks[RKJPEGD_NUM_CLOCKS]; + struct reset_control *resets; + void __iomem *regs; + struct mutex vdev_lock; /* serialises ioctls */ + struct delayed_work watchdog_work; +}; + +/** + * struct rkjpegd_ctx - one open file handle + * + * @fh: V4L2 file handle. + * @dev: device this context belongs to. + * @src_fmt: coded format on the output queue. + * @dst_fmt: raw format on the capture queue. + * @crop: visible part of a capture buffer. The decoder + * writes whole macroblocks, so a frame whose + * height is not a multiple of one is padded. + * @sequence_cap: capture buffer sequence counter. + * @sequence_out: output buffer sequence counter. + * @source_change: a resolution change was reported and the + * capture queue has not been set up again yet. + * @initial_source_change: the first parsed header must report a change + * even when it matches the negotiated format. + * @table_base: quantisation and Huffman table side buffer, + * rebuilt from the frame header on every run. + * + * @source_change and @initial_source_change, and the @dst_fmt and @crop they + * renegotiate, are written from ioctl and videobuf2 callbacks, which hold + * @rkjpegd_dev.vdev_lock, and from rkjpegd_device_run(), which does not: the + * mem2mem core runs it from its job workqueue as well as from a queue + * callback. Nothing serialises the two, but @dst_fmt and @crop are complete + * before the source change event is queued, and an application is told to + * query them only once it has dequeued that event, so one that follows the + * decoder interface cannot observe them half updated. + */ +struct rkjpegd_ctx { + struct v4l2_fh fh; + struct rkjpegd_dev *dev; + struct v4l2_pix_format_mplane src_fmt; + struct v4l2_pix_format_mplane dst_fmt; + struct v4l2_rect crop; + u32 sequence_cap; + u32 sequence_out; + bool source_change; + bool initial_source_change; + struct rkjpegd_aux_buf table_base; +}; + +static inline struct rkjpegd_ctx *file_to_rkjpegd_ctx(struct file *filp) +{ + return container_of(file_to_v4l2_fh(filp), struct rkjpegd_ctx, fh); +} + +static inline void rkjpegd_write_relaxed(struct rkjpegd_dev *jpegd, + u32 val, u32 reg) +{ + writel_relaxed(val, jpegd->regs + reg); +} + +static inline void rkjpegd_write(struct rkjpegd_dev *jpegd, u32 val, u32 reg) +{ + writel(val, jpegd->regs + reg); +} + +static inline u32 rkjpegd_read(struct rkjpegd_dev *jpegd, u32 reg) +{ + return readl(jpegd->regs + reg); +} + +/* + * Every address register is a single 32 bit one and there is no high order + * companion anywhere in the map, so rkjpegd_probe() caps the DMA masks at 32 + * bits and an address that reaches here always fits. + */ +static inline void rkjpegd_write_addr(struct rkjpegd_dev *jpegd, u32 reg, + dma_addr_t addr) +{ + rkjpegd_write(jpegd, lower_32_bits(addr), reg); +} + +static const struct v4l2_event rkjpegd_eos_event = { + .type = V4L2_EVENT_EOS, +}; + +static const struct v4l2_event rkjpegd_src_change_event = { + .type = V4L2_EVENT_SOURCE_CHANGE, + .u.src_change.changes = V4L2_EVENT_SRC_CH_RESOLUTION, +}; + +/* + * Format helpers + * + * The coded queue carries JPEG, the raw queue NV12. The hardware output + * converter has one destination layout, so there is nothing to enumerate on + * either side beyond that single pair. + */ + +/* + * Colorimetry is a property of the picture, and a JPEG frame header carries + * nothing that describes it. Seed both queues with what JFIF implies and + * from then on propagate whatever userspace sets on one queue to the other, + * rather than overriding it: an application that knows the picture is BT.709 + * has no other way to say so, and one that has its format overridden takes it + * as the decoder refusing the stream. + */ +static void rkjpegd_set_default_colorimetry(struct v4l2_pix_format_mplane *pix_mp) +{ + pix_mp->colorspace = V4L2_COLORSPACE_JPEG; + pix_mp->ycbcr_enc = V4L2_YCBCR_ENC_DEFAULT; + pix_mp->quantization = V4L2_QUANTIZATION_DEFAULT; + pix_mp->xfer_func = V4L2_XFER_FUNC_DEFAULT; +} + +static void rkjpegd_propagate_colorimetry(struct v4l2_pix_format_mplane *to, + const struct v4l2_pix_format_mplane *from) +{ + to->colorspace = from->colorspace; + to->ycbcr_enc = from->ycbcr_enc; + to->quantization = from->quantization; + to->xfer_func = from->xfer_func; +} + +static void rkjpegd_fill_raw_fmt(struct v4l2_pix_format_mplane *pix_mp, + u32 width, u32 height) +{ + pix_mp->pixelformat = V4L2_PIX_FMT_NV12; + pix_mp->width = width; + pix_mp->height = height; + pix_mp->field = V4L2_FIELD_NONE; + pix_mp->num_planes = 1; + pix_mp->plane_fmt[0].bytesperline = width; + pix_mp->plane_fmt[0].sizeimage = width * height * 3 / 2; + memset(pix_mp->plane_fmt[0].reserved, 0, + sizeof(pix_mp->plane_fmt[0].reserved)); + memset(pix_mp->reserved, 0, sizeof(pix_mp->reserved)); +} + +static void rkjpegd_fill_coded_fmt(struct v4l2_pix_format_mplane *pix_mp, + u32 width, u32 height, u32 sizeimage) +{ + pix_mp->pixelformat = V4L2_PIX_FMT_JPEG; + pix_mp->width = width; + pix_mp->height = height; + pix_mp->field = V4L2_FIELD_NONE; + pix_mp->num_planes = 1; + pix_mp->plane_fmt[0].bytesperline = 0; + + /* + * A compressed frame has no size the driver could derive, so let + * userspace ask for one and only fall back to an upper bound when it + * does not. A frame that then does not fit is caught when its header + * is parsed, where the missing end of image marker gives it away. + */ + if (!sizeimage) + sizeimage = width * height * RKJPEGD_CODED_MAX_DEPTH; + pix_mp->plane_fmt[0].sizeimage = sizeimage; + + memset(pix_mp->plane_fmt[0].reserved, 0, + sizeof(pix_mp->plane_fmt[0].reserved)); + memset(pix_mp->reserved, 0, sizeof(pix_mp->reserved)); +} + +static void rkjpegd_reset_fmts(struct rkjpegd_ctx *ctx) +{ + u32 width = ALIGN(RKJPEGD_MIN_WIDTH, RKJPEGD_RAW_STEP); + u32 height = ALIGN(RKJPEGD_MIN_HEIGHT, RKJPEGD_RAW_STEP); + + rkjpegd_fill_coded_fmt(&ctx->src_fmt, width, height, 0); + rkjpegd_fill_raw_fmt(&ctx->dst_fmt, width, height); + rkjpegd_set_default_colorimetry(&ctx->src_fmt); + rkjpegd_set_default_colorimetry(&ctx->dst_fmt); + ctx->crop.left = 0; + ctx->crop.top = 0; + ctx->crop.width = width; + ctx->crop.height = height; +} + +/* + * ioctl operations + */ + +static int rkjpegd_querycap(struct file *file, void *priv, + struct v4l2_capability *cap) +{ + strscpy(cap->driver, RKJPEGD_NAME, sizeof(cap->driver)); + strscpy(cap->card, RKJPEGD_NAME, sizeof(cap->card)); + + return 0; +} + +static int rkjpegd_enum_fmt_vid_cap(struct file *file, void *priv, + struct v4l2_fmtdesc *f) +{ + if (f->index) + return -EINVAL; + + f->pixelformat = V4L2_PIX_FMT_NV12; + + return 0; +} + +static int rkjpegd_enum_fmt_vid_out(struct file *file, void *priv, + struct v4l2_fmtdesc *f) +{ + if (f->index) + return -EINVAL; + + f->pixelformat = V4L2_PIX_FMT_JPEG; + + /* + * The frame dimensions come out of the bitstream rather than out of + * the coded format, which is what V4L2_EVENT_SOURCE_CHANGE reports. + * Userspace has no other way to find out that the event is worth + * waiting for, and one that does not wait sets the capture queue up + * from whatever the coded format happened to say. + */ + f->flags |= V4L2_FMT_FLAG_DYN_RESOLUTION; + + return 0; +} + +static int rkjpegd_enum_framesizes(struct file *file, void *priv, + struct v4l2_frmsizeenum *fsize) +{ + if (fsize->index) + return -EINVAL; + + /* + * Only the coded format has a size userspace gets to pick. The + * decoded one follows the bitstream, so there is nothing to enumerate + * for it, and saying otherwise is actively harmful: a step here + * describes the sizes that may be asked for, and userspace turns it + * into a set the visible rectangle then has to be a member of. The + * capture buffer is padded to a macroblock, the picture in it is not, + * so a 1080 line frame would fall outside a set built on a step of 16 + * and could not be negotiated at all. + */ + if (fsize->pixel_format != V4L2_PIX_FMT_JPEG) + return -ENOTTY; + + fsize->type = V4L2_FRMSIZE_TYPE_STEPWISE; + fsize->stepwise.min_width = RKJPEGD_MIN_WIDTH; + fsize->stepwise.max_width = RKJPEGD_MAX_SIZE; + fsize->stepwise.step_width = RKJPEGD_CODED_STEP; + fsize->stepwise.min_height = RKJPEGD_MIN_HEIGHT; + fsize->stepwise.max_height = RKJPEGD_MAX_SIZE; + fsize->stepwise.step_height = RKJPEGD_CODED_STEP; + + return 0; +} + +static int rkjpegd_g_fmt_vid_cap(struct file *file, void *priv, + struct v4l2_format *f) +{ + f->fmt.pix_mp = file_to_rkjpegd_ctx(file)->dst_fmt; + + return 0; +} + +static int rkjpegd_g_fmt_vid_out(struct file *file, void *priv, + struct v4l2_format *f) +{ + f->fmt.pix_mp = file_to_rkjpegd_ctx(file)->src_fmt; + + return 0; +} + +/* + * The capture resolution follows the bitstream, not userspace: it is set from + * the frame header when a source change is reported and only read back here. + */ +static int rkjpegd_try_fmt_vid_cap(struct file *file, void *priv, + struct v4l2_format *f) +{ + struct rkjpegd_ctx *ctx = file_to_rkjpegd_ctx(file); + + rkjpegd_fill_raw_fmt(&f->fmt.pix_mp, ctx->dst_fmt.width, + ctx->dst_fmt.height); + + return 0; +} + +static int rkjpegd_try_fmt_vid_out(struct file *file, void *priv, + struct v4l2_format *f) +{ + struct v4l2_pix_format_mplane *pix_mp = &f->fmt.pix_mp; + u32 sizeimage = pix_mp->num_planes == 1 ? + pix_mp->plane_fmt[0].sizeimage : 0; + + v4l_bound_align_image(&pix_mp->width, + RKJPEGD_MIN_WIDTH, RKJPEGD_MAX_SIZE, + ilog2(RKJPEGD_CODED_STEP), + &pix_mp->height, + RKJPEGD_MIN_HEIGHT, RKJPEGD_MAX_SIZE, + ilog2(RKJPEGD_CODED_STEP), 0); + + rkjpegd_fill_coded_fmt(pix_mp, pix_mp->width, pix_mp->height, sizeimage); + + return 0; +} + +static int rkjpegd_s_fmt_vid_cap(struct file *file, void *priv, + struct v4l2_format *f) +{ + struct rkjpegd_ctx *ctx = file_to_rkjpegd_ctx(file); + struct vb2_queue *vq = v4l2_m2m_get_dst_vq(ctx->fh.m2m_ctx); + + if (vb2_is_busy(vq)) + return -EBUSY; + + rkjpegd_try_fmt_vid_cap(file, priv, f); + ctx->dst_fmt = f->fmt.pix_mp; + rkjpegd_propagate_colorimetry(&ctx->src_fmt, &ctx->dst_fmt); + + return 0; +} + +static int rkjpegd_s_fmt_vid_out(struct file *file, void *priv, + struct v4l2_format *f) +{ + struct rkjpegd_ctx *ctx = file_to_rkjpegd_ctx(file); + struct vb2_queue *vq = v4l2_m2m_get_src_vq(ctx->fh.m2m_ctx); + int ret; + + if (vb2_is_busy(vq)) + return -EBUSY; + + ret = rkjpegd_try_fmt_vid_out(file, priv, f); + if (ret) + return ret; + + ctx->src_fmt = f->fmt.pix_mp; + + /* + * Picking a coded format invalidates whatever was negotiated on the + * capture queue. Seed it from the coded resolution so a userspace + * that already knows the frame size can allocate without waiting for + * the first source change. + */ + rkjpegd_fill_raw_fmt(&ctx->dst_fmt, + ALIGN(ctx->src_fmt.width, RKJPEGD_RAW_STEP), + ALIGN(ctx->src_fmt.height, RKJPEGD_RAW_STEP)); + rkjpegd_propagate_colorimetry(&ctx->dst_fmt, &ctx->src_fmt); + ctx->crop.left = 0; + ctx->crop.top = 0; + ctx->crop.width = ctx->src_fmt.width; + ctx->crop.height = ctx->src_fmt.height; + + return 0; +} + +static int rkjpegd_g_selection(struct file *file, void *priv, + struct v4l2_selection *s) +{ + struct rkjpegd_ctx *ctx = file_to_rkjpegd_ctx(file); + + if (s->type != V4L2_BUF_TYPE_VIDEO_CAPTURE && + s->type != V4L2_BUF_TYPE_VIDEO_CAPTURE_MPLANE) + return -EINVAL; + + switch (s->target) { + case V4L2_SEL_TGT_COMPOSE: + case V4L2_SEL_TGT_COMPOSE_DEFAULT: + s->r = ctx->crop; + break; + case V4L2_SEL_TGT_COMPOSE_BOUNDS: + case V4L2_SEL_TGT_COMPOSE_PADDED: + s->r.left = 0; + s->r.top = 0; + s->r.width = ctx->dst_fmt.width; + s->r.height = ctx->dst_fmt.height; + break; + default: + return -EINVAL; + } + + return 0; +} + +static int rkjpegd_subscribe_event(struct v4l2_fh *fh, + const struct v4l2_event_subscription *sub) +{ + switch (sub->type) { + case V4L2_EVENT_EOS: + return v4l2_event_subscribe(fh, sub, 0, NULL); + case V4L2_EVENT_SOURCE_CHANGE: + return v4l2_src_change_event_subscribe(fh, sub); + default: + /* The decoder takes no controls, so there is nothing else. */ + return -EINVAL; + } +} + +static void rkjpegd_set_last_buffer(struct rkjpegd_ctx *ctx) +{ + struct vb2_v4l2_buffer *next_dst_buf; + + next_dst_buf = v4l2_m2m_dst_buf_remove(ctx->fh.m2m_ctx); + if (!next_dst_buf) { + ctx->fh.m2m_ctx->is_draining = true; + ctx->fh.m2m_ctx->next_buf_last = true; + return; + } + + v4l2_m2m_last_buffer_done(ctx->fh.m2m_ctx, next_dst_buf); +} + +static int rkjpegd_decoder_cmd(struct file *file, void *priv, + struct v4l2_decoder_cmd *cmd) +{ + struct rkjpegd_ctx *ctx = file_to_rkjpegd_ctx(file); + int ret; + + ret = v4l2_m2m_ioctl_try_decoder_cmd(file, priv, cmd); + if (ret < 0) + return ret; + + if (!vb2_is_streaming(v4l2_m2m_get_src_vq(ctx->fh.m2m_ctx))) + return 0; + + ret = v4l2_m2m_ioctl_decoder_cmd(file, priv, cmd); + if (ret < 0) + return ret; + + if (cmd->cmd == V4L2_DEC_CMD_STOP && + v4l2_m2m_has_stopped(ctx->fh.m2m_ctx)) + v4l2_event_queue_fh(&ctx->fh, &rkjpegd_eos_event); + + if (cmd->cmd == V4L2_DEC_CMD_START && + v4l2_m2m_has_stopped(ctx->fh.m2m_ctx)) + vb2_clear_last_buffer_dequeued(&ctx->fh.m2m_ctx->cap_q_ctx.q); + + return 0; +} + +static const struct v4l2_ioctl_ops rkjpegd_ioctl_ops = { + .vidioc_querycap = rkjpegd_querycap, + .vidioc_enum_framesizes = rkjpegd_enum_framesizes, + + .vidioc_enum_fmt_vid_cap = rkjpegd_enum_fmt_vid_cap, + .vidioc_g_fmt_vid_cap_mplane = rkjpegd_g_fmt_vid_cap, + .vidioc_try_fmt_vid_cap_mplane = rkjpegd_try_fmt_vid_cap, + .vidioc_s_fmt_vid_cap_mplane = rkjpegd_s_fmt_vid_cap, + + .vidioc_enum_fmt_vid_out = rkjpegd_enum_fmt_vid_out, + .vidioc_g_fmt_vid_out_mplane = rkjpegd_g_fmt_vid_out, + .vidioc_try_fmt_vid_out_mplane = rkjpegd_try_fmt_vid_out, + .vidioc_s_fmt_vid_out_mplane = rkjpegd_s_fmt_vid_out, + + .vidioc_g_selection = rkjpegd_g_selection, + + .vidioc_reqbufs = v4l2_m2m_ioctl_reqbufs, + .vidioc_querybuf = v4l2_m2m_ioctl_querybuf, + .vidioc_qbuf = v4l2_m2m_ioctl_qbuf, + .vidioc_dqbuf = v4l2_m2m_ioctl_dqbuf, + .vidioc_prepare_buf = v4l2_m2m_ioctl_prepare_buf, + .vidioc_create_bufs = v4l2_m2m_ioctl_create_bufs, + .vidioc_expbuf = v4l2_m2m_ioctl_expbuf, + .vidioc_remove_bufs = v4l2_m2m_ioctl_remove_bufs, + + .vidioc_streamon = v4l2_m2m_ioctl_streamon, + .vidioc_streamoff = v4l2_m2m_ioctl_streamoff, + + .vidioc_try_decoder_cmd = v4l2_m2m_ioctl_try_decoder_cmd, + .vidioc_decoder_cmd = rkjpegd_decoder_cmd, + + .vidioc_subscribe_event = rkjpegd_subscribe_event, + .vidioc_unsubscribe_event = v4l2_event_unsubscribe, +}; + +/* + * Job handling + */ + +static void rkjpegd_arm_watchdog(struct rkjpegd_dev *jpegd) +{ + schedule_delayed_work(&jpegd->watchdog_work, + msecs_to_jiffies(RKJPEGD_TIMEOUT_MS)); +} + +static void rkjpegd_job_finish_no_pm(struct rkjpegd_ctx *ctx, + enum vb2_buffer_state state) +{ + struct vb2_v4l2_buffer *src, *dst; + + src = v4l2_m2m_next_src_buf(ctx->fh.m2m_ctx); + dst = v4l2_m2m_next_dst_buf(ctx->fh.m2m_ctx); + if (WARN_ON(!src) || WARN_ON(!dst)) + return; + + src->sequence = ctx->sequence_out++; + dst->sequence = ctx->sequence_cap++; + + if (v4l2_m2m_is_last_draining_src_buf(ctx->fh.m2m_ctx, src)) { + dst->flags |= V4L2_BUF_FLAG_LAST; + v4l2_event_queue_fh(&ctx->fh, &rkjpegd_eos_event); + v4l2_m2m_mark_stopped(ctx->fh.m2m_ctx); + } + + v4l2_m2m_buf_done_and_job_finish(ctx->dev->m2m_dev, ctx->fh.m2m_ctx, + state); +} + +static void rkjpegd_job_finish(struct rkjpegd_ctx *ctx, + enum vb2_buffer_state state) +{ + struct rkjpegd_dev *jpegd = ctx->dev; + + clk_bulk_disable(RKJPEGD_NUM_CLOCKS, jpegd->clocks); + pm_runtime_put_autosuspend(jpegd->dev); + + rkjpegd_job_finish_no_pm(ctx, state); +} + +static void rkjpegd_irq_done(struct rkjpegd_dev *jpegd, + enum vb2_buffer_state state) +{ + struct rkjpegd_ctx *ctx = v4l2_m2m_get_curr_priv(jpegd->m2m_dev); + + if (!ctx) + return; + + /* + * A false return means the watchdog already expired and is running, + * so it owns finishing this job. + */ + if (cancel_delayed_work(&jpegd->watchdog_work)) + rkjpegd_job_finish(ctx, state); +} + +/* + * VDPU720 register programming + */ + +/** + * vdpu720_jpeg_mode() - map JPEG sampling factors to the hardware mode + * @frame: parsed JPEG frame header + * + * The sampling factor tuple is determined by the luma channel's factors + * relative to the maximum in the frame. Standard JFIF layouts only, + * anything else returns -EINVAL: the mode also picks the MCU height and + * therefore PIC_H, so guessing one would decode into a wrong image. + * + * Return: a VDPU720_JPEG_MODE_* value, or -EINVAL for a sampling layout + * the hardware cannot be told about. + */ +static int vdpu720_jpeg_mode(const struct v4l2_jpeg_frame_header *frame) +{ + u8 h0, v0; + + if (frame->num_components == 1) + return VDPU720_JPEG_MODE_YUV400; + + /* Component 0 always carries luma in JFIF */ + h0 = frame->component[0].horizontal_sampling_factor; + v0 = frame->component[0].vertical_sampling_factor; + + if (h0 == 1 && v0 == 1) + return VDPU720_JPEG_MODE_YUV444; + if (h0 == 2 && v0 == 1) + return VDPU720_JPEG_MODE_YUV422; + if (h0 == 2 && v0 == 2) + return VDPU720_JPEG_MODE_YUV420; + if (h0 == 4 && v0 == 1) + return VDPU720_JPEG_MODE_YUV411; + if (h0 == 1 && v0 == 2) + return VDPU720_JPEG_MODE_YUV440; + + return -EINVAL; +} + +/** + * vdpu720_nb_htbl_sets() - number of Huffman table sets the hardware reads + * @num_components: number of components in the scan + * + * One for a grayscale frame, two for a colour one. vdpu720_write_htbl() + * fills that many sets and vdpu720_fill_regs() sizes HTBL_SEL and the length + * registers from the same number, so the two cannot drift apart. + * + * Return: the number of table sets to write. + */ +static unsigned int vdpu720_nb_htbl_sets(unsigned int num_components) +{ + return num_components == 1 ? 1 : VDPU720_NB_HTBL_SETS; +} + +/** + * vdpu720_write_qtbl() - write all Q-tables into the DMA side buffer + * @ctx: context whose side buffer receives the tables + * @hdr: parsed JPEG header the tables are taken from + * + * Tables are stored sequentially, one per component in component order. + * Each entry is widened to u16 and reordered from JPEG zig-zag scan to + * natural raster-scan order, matching the hardware expectation. + * + * Return: 0 on success, -EINVAL if the frame refers to a quantization + * table it does not carry. + */ +static int vdpu720_write_qtbl(struct rkjpegd_ctx *ctx, + const struct v4l2_jpeg_header *hdr) +{ + struct rkjpegd_dev *jpegd = ctx->dev; + u16 *base = ctx->table_base.cpu; + unsigned int k, i; + + for (k = 0; k < hdr->frame.num_components; k++) { + u8 tq_id = hdr->frame.component[k].quantization_table_selector; + u8 qtbl[VDPU720_QTBL_ENTRIES]; + u16 *dst; + + /* + * v4l2_jpeg_parse_header() filled quantization_tables[] by + * destination selector (packed DQT segments handled); .start + * points at the 64 Qk values, already past the Pq|Tq byte. + */ + if (tq_id > 3 || !hdr->quantization_tables[tq_id].start) { + dev_err_ratelimited(jpegd->dev, + "Q-table %u not found for component %u\n", + tq_id, k); + return -EINVAL; + } + + /* + * Bulk-copy the Q-table out of the uncached source buffer once; + * the per-element zigzag reads below would otherwise each be a + * separate uncached bus transaction. + */ + memcpy(qtbl, hdr->quantization_tables[tq_id].start, sizeof(qtbl)); + dst = base + k * VDPU720_QTBL_ENTRIES; + + /* + * Reorder zigzag -> raster scan. + * v4l2_jpeg_zigzag_scan_index[z] = raster position of zigzag + * element z. JPEG Q-tables are stored in zigzag order; the + * hardware expects them in natural raster (row-major) order. + */ + for (i = 0; i < VDPU720_QTBL_ENTRIES; i++) + dst[v4l2_jpeg_zigzag_scan_index[i]] = (u16)qtbl[i]; + } + + return 0; +} + +/** + * vdpu720_compute_mincode() - build the minimum Huffman code arrays + * @bits: BITS[16], the number of codes of each length 1..16 + * @min_code: output, minimum code value per length (16 entries) + * @acc_addr: output, accumulated symbol-table address per length (16 entries) + * + * Derives the two arrays the hardware needs for one Huffman table, DC or AC, + * from that table's BITS array. Algorithm ported verbatim from + * jpegd_vpu7xx_write_htbl(). + */ +static void vdpu720_compute_mincode(const u8 *bits, u16 *min_code, u16 *acc_addr) +{ + u16 code = 0, addr = 0; + unsigned int j; + + for (j = 0; j < 16; j++) { + u16 len = bits[j]; + + if (len == 0 && j > 0) { + if (code > ((u16)(min_code[j - 1]) << 1)) + min_code[j] = code; + else + min_code[j] = (u16)(min_code[j - 1]) << 1; + } else { + min_code[j] = code; + } + + code += len; + addr += len; + acc_addr[j] = addr; + code <<= 1; + } + + /* Sentinel: set min_code[0] to the last valid code + count */ + if (bits[15]) + min_code[0] = min_code[15] + bits[15] - 1; + else + min_code[0] = min_code[15]; +} + +/** + * vdpu720_write_htbl() - fill the Huffman mincode and value sub-buffers + * @ctx: context whose side buffer receives the tables + * @hdr: parsed JPEG header the tables are taken from + * + * One set is written per vdpu720_nb_htbl_sets(), fed by the scan component + * that uses it: the first by the luma component and the second by the first + * chroma one. The hardware has no room for a third set and no per component + * selector register, so a frame whose two chroma components disagree on their + * tables cannot be described to it and is refused rather than decoded with + * the wrong table for the last component. + * + * Per-set layout in the mincode buffer: + * 16 x u16 DC min-codes + * 8 x u16 DC accumulated addresses (packed pairs) + * 16 x u16 AC min-codes + * 8 x u16 AC accumulated addresses (packed pairs) + * + * Per-set layout in the value buffer (192 bytes): + * 16 bytes DC code values + * 176 bytes AC code values + * + * Return: 0 on success, -EINVAL if the scan refers to a Huffman table it + * does not carry or needs more table sets than the hardware has. + */ +static int vdpu720_write_htbl(struct rkjpegd_ctx *ctx, + const struct v4l2_jpeg_header *hdr) +{ + struct rkjpegd_dev *jpegd = ctx->dev; + const struct v4l2_jpeg_scan_header *scan = hdr->scan; + u8 *tbl_base = ctx->table_base.cpu; + u16 *p_mincode = (u16 *)(tbl_base + VDPU720_HMINCODE_OFF); + u8 *p_value = tbl_base + VDPU720_HVALUE_OFF; + unsigned int nb_sets = vdpu720_nb_htbl_sets(scan->num_components); + unsigned int k, i; + + /* The last set is shared by every remaining component */ + for (k = nb_sets; k < scan->num_components; k++) { + if (scan->component[k].dc_entropy_coding_table_selector != + scan->component[nb_sets - 1].dc_entropy_coding_table_selector || + scan->component[k].ac_entropy_coding_table_selector != + scan->component[nb_sets - 1].ac_entropy_coding_table_selector) { + dev_err_ratelimited(jpegd->dev, + "JPEG component %u uses other Huffman tables than component %u\n", + k, nb_sets - 1); + return -EINVAL; + } + } + + for (k = 0; k < nb_sets; k++) { + u8 dc_sel = scan->component[k].dc_entropy_coding_table_selector; + u8 ac_sel = scan->component[k].ac_entropy_coding_table_selector; + u8 dc_bits[16], ac_bits[16]; + const u8 *dc_src, *ac_src, *dc_vals, *ac_vals; + unsigned int dc_huffval_len, ac_huffval_len; + u16 min_dc[16], acc_dc[16]; + u16 min_ac[16], acc_ac[16]; + + /* + * v4l2_jpeg_parse_header() filled huffman_tables[] indexed by + * (Tc << 1) | Th - Tc=0/1 (DC/AC) class, Th=0/1 (luma/chroma) + * id (packed DHT segments handled). .start points at BITS[16], + * already past the Tc|Th byte. + */ + if (dc_sel > 1 || ac_sel > 1 || + !hdr->huffman_tables[dc_sel].start || + !hdr->huffman_tables[2 | ac_sel].start) { + dev_err_ratelimited(jpegd->dev, + "H-table not found for component %u (dc=%u ac=%u)\n", + k, dc_sel, ac_sel); + return -EINVAL; + } + + /* + * Layout at .start: [BITS 16B] [HUFFVAL sum(BITS)B]. Derive the + * HUFFVAL length from BITS (a packed segment's length would + * over-count when several tables share it). + */ + dc_src = hdr->huffman_tables[dc_sel].start; + ac_src = hdr->huffman_tables[2 | ac_sel].start; + + /* + * Bulk-copy the two BITS arrays out of the uncached source; they + * are otherwise walked byte-by-byte twice (the length sum here and + * again in vdpu720_compute_mincode()). The HUFFVAL blocks stay in + * the source and are copied out in one memcpy() further down. + */ + memcpy(dc_bits, dc_src, sizeof(dc_bits)); + memcpy(ac_bits, ac_src, sizeof(ac_bits)); + dc_vals = dc_src + 16; + ac_vals = ac_src + 16; + + dc_huffval_len = 0; + for (i = 0; i < 16; i++) + dc_huffval_len += dc_bits[i]; + ac_huffval_len = 0; + for (i = 0; i < 16; i++) + ac_huffval_len += ac_bits[i]; + + /* + * The value block holds VDPU720_DC_VALUES_MAX DC and + * VDPU720_AC_VALUES_MAX AC symbols, and the accumulated + * addresses below are packed two per u16, so they have to stay + * within a byte as well. A table the hardware cannot hold + * would otherwise be truncated into it silently and decode to + * a wrong image. Both limits are covered by this check, the + * value block is the tighter of the two. + */ + if (dc_huffval_len > VDPU720_DC_VALUES_MAX || + ac_huffval_len > VDPU720_AC_VALUES_MAX) { + dev_err_ratelimited(jpegd->dev, + "JPEG Huffman table too large for component %u (dc=%u ac=%u)\n", + k, dc_huffval_len, ac_huffval_len); + return -EINVAL; + } + + vdpu720_compute_mincode(dc_bits, min_dc, acc_dc); + vdpu720_compute_mincode(ac_bits, min_ac, acc_ac); + + for (i = 0; i < 16; i++) + *p_mincode++ = min_dc[i]; + for (i = 0; i < 8; i++) + *p_mincode++ = (u16)acc_dc[2 * i] | + ((u16)acc_dc[2 * i + 1] << 8); + for (i = 0; i < 16; i++) + *p_mincode++ = min_ac[i]; + for (i = 0; i < 8; i++) + *p_mincode++ = (u16)acc_ac[2 * i] | + ((u16)acc_ac[2 * i + 1] << 8); + + /* Zero-pad the value block, then fill DC then AC values. */ + memset(p_value, 0, VDPU720_HVALUE_SET_SIZE); + memcpy(p_value, dc_vals, dc_huffval_len); + memcpy(p_value + VDPU720_DC_VALUES_MAX, ac_vals, ac_huffval_len); + p_value += VDPU720_HVALUE_SET_SIZE; + } + + return 0; +} + +/** + * vdpu720_fill_regs() - program the decoder registers for one frame + * @ctx: context the job belongs to + * @hdr: parsed JPEG header + * @tbl_dma: DMA address of the Q/Huffman table side buffer + * @strm_dma: DMA address of the entropy stream, 16-byte aligned + * @strm_start_byte: offset of the first stream byte within that word + * @strm_len_blks: stream length in 16-byte blocks, minus one + * @out_dma: DMA address of the destination NV12 buffer + * + * Return: 0 on success, negative errno for a frame the hardware cannot be + * programmed for. + */ +static int vdpu720_fill_regs(struct rkjpegd_ctx *ctx, + const struct v4l2_jpeg_header *hdr, + dma_addr_t tbl_dma, + dma_addr_t strm_dma, u32 strm_start_byte, + u32 strm_len_blks, + dma_addr_t out_dma) +{ + struct rkjpegd_dev *jpegd = ctx->dev; + /* + * Use negotiated buffer dimensions for stride/vstride so the NV12 UV + * plane lands at the correct offset in the allocated buffer. + * + * For PIC_SIZE, use the MCU-boundary-aligned height rather than the + * raw JPEG header height. The hardware computes its vertical MCU count + * as floor(PIC_H / mcu_height). For YUV420/YUV440 (mcu_height=16), + * a 1080-pixel-high image yields floor(1080/16)=67 MCUs, but the JPEG + * encoder always writes ceil(1080/16)=68 complete MCUs (1088 rows of + * entropy data). Using PIC_H=1080 therefore causes the hardware to + * stop after row 1072, leaving the bottom 8 rows unwritten and + * potentially triggering decode errors from unread bitstream bytes. + * + * Using ALIGN(jpeg_height, mcu_height) is always safe: the encoder + * writes exactly ceil(height/mcu_height) MCUs, so the aligned height + * matches the actual entropy data in the bitstream. For modes where + * jpeg_height is already on an MCU boundary (all 8-pixel-MCU modes at + * standard resolutions), ALIGN() is a no-op. + */ + u32 jpeg_width = hdr->frame.width; + u32 jpeg_height = hdr->frame.height; + u32 buf_width = ctx->dst_fmt.width; + u32 buf_height = ctx->dst_fmt.height; + u32 w_align = ALIGN(buf_width, 16); + u32 y_stride = w_align >> 4; /* units of 16 pixels */ + u32 y_vstride = y_stride * buf_height; /* sets the UV plane offset */ + u32 nb_comp = hdr->frame.num_components; + /* + * qtbl_sel = number of Q-table entries written to the side buffer, + * one per component. Using the DQT segment count is wrong when a + * camera packs all Q-tables into a single segment while there are + * three components: the hardware would only read one Q-table, leaving + * Cb/Cr with garbage. + */ + u32 qtbl_sel = nb_comp; + /* + * H-table sets: one for grayscale (luma only), two for colour (luma + + * chroma). The three table lengths below are computed from these two + * counts and the side buffer layout, so what is programmed is what + * vdpu720_write_qtbl() and vdpu720_write_htbl() actually wrote. + */ + u32 htbl_sel = vdpu720_nb_htbl_sets(nb_comp); + u32 mcu_width, mcu_height, jpeg_height_aligned; + u32 qtbl_len, hmin_len, hval_len; + int jpeg_mode; + u32 reg; + + jpeg_mode = vdpu720_jpeg_mode(&hdr->frame); + if (jpeg_mode < 0) { + dev_err_ratelimited(jpegd->dev, + "unsupported JPEG sampling factors %ux%u\n", + hdr->frame.component[0].horizontal_sampling_factor, + hdr->frame.component[0].vertical_sampling_factor); + return jpeg_mode; + } + + /* + * MCU size follows the luma sampling factors, MCU_W = h0 * 8 and + * MCU_H = v0 * 8. YUV420 and YUV440 subsample the luma vertically by + * 2, giving MCU_H = 16; YUV420 and YUV422 subsample it horizontally + * by 2 and YUV411 by 4, giving MCU_W = 16 and 32. The rest is 8. + */ + mcu_height = (jpeg_mode == VDPU720_JPEG_MODE_YUV420 || + jpeg_mode == VDPU720_JPEG_MODE_YUV440) ? 16 : 8; + mcu_width = (jpeg_mode == VDPU720_JPEG_MODE_YUV411) ? 32 : + (jpeg_mode == VDPU720_JPEG_MODE_YUV420 || + jpeg_mode == VDPU720_JPEG_MODE_YUV422) ? 16 : 8; + jpeg_height_aligned = ALIGN(jpeg_height, mcu_height); + + /* + * The picture dimensions below are taken from the bitstream while the + * strides are taken from the negotiated capture format. A frame that + * is larger than what was negotiated would make the decoder write + * beyond the capture buffer, so refuse it rather than program the + * hardware with the two sets of numbers mixed. + */ + if (jpeg_width > buf_width || jpeg_height_aligned > buf_height) { + dev_err_ratelimited(jpegd->dev, + "JPEG %ux%u does not fit the negotiated %ux%u\n", + jpeg_width, jpeg_height_aligned, + buf_width, buf_height); + return -EINVAL; + } + + /* + * REG2: system config - always output NV12. + * + * FILL_DOWN_E belongs to the NV12 conversion rather than to a + * particular pair of heights: the output chroma is vertically + * subsampled, so the hardware has to complete the bottom of the + * picture. Two cases need it: + * + * a) jpeg_height is not on an MCU boundary (e.g. YUV420 1080p: + * jpeg_height=1080, jpeg_height_aligned=1088, buf_height=1088). + * The hardware needs FILL_DOWN_E to complete the last MCU row's + * chroma reconstruction. Without it the bottom rows are corrupt + * even though PIC_H is already set to the MCU-aligned height. + * + * b) MCU-aligned height < buf_height (e.g. YUV422 1080p: mcu_h=8, + * jpeg_height_aligned=1080, buf_height=1088). The hardware fills + * rows 1080..1087 by repeating the last valid row. + * + * A height that is a multiple of 16 leaves nothing to fill, but the + * bit is set there as well: the reference driver enables it for every + * NV12 conversion, whatever the picture and buffer heights are. + * + * FILL_RIGHT_E does the same for the right hand edge, but there it + * depends on the mode. The decoder writes whole MCUs, so the last + * MCU column ends at ALIGN(jpeg_width, mcu_width) while the buffer is + * 16 pixel aligned. Only the 8 pixel MCU widths can stop short of + * that and need the columns in between filled; a 16 or 32 pixel MCU + * already reaches at least as far. The reference driver arrives at + * the same set through a per mode test on (width & 0xf) <= 8. + */ + reg = FIELD_PREP(VDPU720_YUV_OUT_FMT, VDPU720_YUV_OUT_FMT_NV12) | + VDPU720_FILL_DOWN_E; + if (ALIGN(jpeg_width, mcu_width) < ALIGN(jpeg_width, 16)) + reg |= VDPU720_FILL_RIGHT_E; + rkjpegd_write_relaxed(jpegd, reg, VDPU720_REG_SYS); + + /* + * REG3: picture dimensions + * + * PIC_W stays at the raw header width while PIC_H is rounded up to + * the MCU boundary. Rounding the width up the same way is not safe: + * mcu_width reaches 32 for YUV411, so ALIGN(jpeg_width, mcu_width) + * can land beyond the 16 pixel aligned buffer width and point the + * decoder past the end of a row. FILL_RIGHT_E above covers the cases + * where the MCU column stops short instead. The reference driver + * programs the raw width here too. + */ + rkjpegd_write_relaxed(jpegd, + FIELD_PREP(VDPU720_PIC_W_M1, jpeg_width - 1) | + FIELD_PREP(VDPU720_PIC_H_M1, jpeg_height_aligned - 1), + VDPU720_REG_PIC_SIZE); + + /* REG4: JPEG format, Q/H table counts, restart interval */ + qtbl_len = VDPU720_TBL_LEN(qtbl_sel * VDPU720_QTBL_COMP_SIZE); + hmin_len = VDPU720_TBL_LEN(htbl_sel * VDPU720_HMINCODE_SET_SIZE); + hval_len = VDPU720_TBL_LEN(htbl_sel * VDPU720_HVALUE_SET_SIZE); + + reg = FIELD_PREP(VDPU720_JPEG_MODE, jpeg_mode) | + FIELD_PREP(VDPU720_PIX_DEPTH, VDPU720_PIX_DEPTH_8) | + FIELD_PREP(VDPU720_QTBL_SEL, qtbl_sel) | + FIELD_PREP(VDPU720_HTBL_SEL, htbl_sel); + if (hdr->restart_interval) { + reg |= VDPU720_DRI_E; + reg |= FIELD_PREP(VDPU720_DRI_MCU_M1, hdr->restart_interval - 1); + } + rkjpegd_write_relaxed(jpegd, reg, VDPU720_REG_PIC_FMT); + + /* REG5: horizontal virtual strides */ + rkjpegd_write_relaxed(jpegd, + FIELD_PREP(VDPU720_Y_HOR_STRIDE, y_stride) | + FIELD_PREP(VDPU720_UV_HOR_STRIDE, y_stride), + VDPU720_REG_HOR_STRIDE); + + /* REG6: total Y-plane size (stride-units * height) */ + rkjpegd_write_relaxed(jpegd, FIELD_PREP(VDPU720_Y_VSTRIDE, y_vstride), + VDPU720_REG_Y_VSTRIDE); + + /* REG7: table lengths + high stride bit */ + reg = FIELD_PREP(VDPU720_QTBL_LEN, qtbl_len) | + FIELD_PREP(VDPU720_HTBL_MINCODE_LEN, hmin_len) | + FIELD_PREP(VDPU720_HTBL_VALUE_LEN, hval_len) | + FIELD_PREP(VDPU720_Y_HOR_STRIDE_H, y_stride >> 16); + rkjpegd_write_relaxed(jpegd, reg, VDPU720_REG_TBL_LEN); + + /* REG8: stream length and start byte */ + rkjpegd_write_relaxed(jpegd, + FIELD_PREP(VDPU720_STRM_START_BYTE, strm_start_byte) | + FIELD_PREP(VDPU720_STRM_LEN, strm_len_blks), + VDPU720_REG_STRM_LEN); + + /* REG9-REG11: Q/H table DMA addresses (side buffer) */ + rkjpegd_write_addr(jpegd, VDPU720_REG_QTBL_BASE, tbl_dma); + rkjpegd_write_addr(jpegd, VDPU720_REG_HTBL_MINCODE, + tbl_dma + VDPU720_HMINCODE_OFF); + rkjpegd_write_addr(jpegd, VDPU720_REG_HTBL_VALUE, + tbl_dma + VDPU720_HVALUE_OFF); + + /* REG12: stream base (16-byte aligned) */ + rkjpegd_write_addr(jpegd, VDPU720_REG_STRM_BASE, strm_dma); + + /* REG13: NV12 output buffer */ + rkjpegd_write_addr(jpegd, VDPU720_REG_OUT_BASE, out_dma); + + /* REG14: stream error handling defaults */ + rkjpegd_write_relaxed(jpegd, VDPU720_STRM_ERR_DFLT, VDPU720_REG_STRM_ERR); + + /* REG16: enable all internal clock gates */ + rkjpegd_write_relaxed(jpegd, VDPU720_CLK_GATE_ALL, VDPU720_REG_CLK_GATE); + + /* REG30: AXI performance counter */ + rkjpegd_write_relaxed(jpegd, + VDPU720_PERF_WORK_E | VDPU720_PERF_CLR_E | + VDPU720_PERF_CNT_TYPE | + FIELD_PREP(VDPU720_PERF_RD_LAT_ID, 0xa), + VDPU720_REG_PERF_CTRL); + + return 0; +} + +/** + * vdpu720_fill_chroma() - write neutral chroma for a grayscale frame + * @ctx: context the job belongs to + * @dst_buf: capture buffer whose chroma plane is filled + * + * The output format converter has no YUV400 path: VDPU720_YUV_OUT_FMT_NV12 + * only covers the subsampled colour modes, and for a single component frame + * the hardware writes the luma plane and leaves the chroma plane untouched. + * + * Fill the plane here, before the hardware is started: once the decode is + * running the interrupt can complete the job and hand the buffer to + * userspace at any time. + * + * Return: 0 on success, -EINVAL if the capture buffer has no kernel mapping. + */ +static int vdpu720_fill_chroma(struct rkjpegd_ctx *ctx, + struct vb2_v4l2_buffer *dst_buf) +{ + struct rkjpegd_dev *jpegd = ctx->dev; + u32 y_size = ctx->dst_fmt.plane_fmt[0].bytesperline * ctx->dst_fmt.height; + u32 size = ctx->dst_fmt.plane_fmt[0].sizeimage; + void *dst_cpu; + + dst_cpu = vb2_plane_vaddr(&dst_buf->vb2_buf, 0); + if (!dst_cpu) { + dev_err_ratelimited(jpegd->dev, + "JPEG capture buffer has no kernel mapping\n"); + return -EINVAL; + } + + memset(dst_cpu + y_size, 0x80, size - y_size); + + return 0; +} + +/** + * rkjpegd_vdpu720_init() - allocate the table side buffer + * @ctx: context to allocate the Q/Huffman table buffer for + * + * Return: 0 on success, -ENOMEM if the allocation failed. + */ +static int rkjpegd_vdpu720_init(struct rkjpegd_ctx *ctx) +{ + struct rkjpegd_dev *jpegd = ctx->dev; + + ctx->table_base.size = VDPU720_TABLE_BUF_SIZE; + ctx->table_base.cpu = dma_alloc_noncoherent(jpegd->dev, + ctx->table_base.size, + &ctx->table_base.dma, + DMA_TO_DEVICE, GFP_KERNEL); + if (!ctx->table_base.cpu) + return -ENOMEM; + + return 0; +} + +/** + * rkjpegd_vdpu720_exit() - free the table side buffer + * @ctx: context the buffer belongs to + */ +static void rkjpegd_vdpu720_exit(struct rkjpegd_ctx *ctx) +{ + struct rkjpegd_dev *jpegd = ctx->dev; + + dma_free_noncoherent(jpegd->dev, ctx->table_base.size, + ctx->table_base.cpu, ctx->table_base.dma, + DMA_TO_DEVICE); +} + +/** + * rkjpegd_vdpu720_run() - fill the side buffer, program the registers and + * start the hardware + * @ctx: context holding the queues and the side buffer + * + * The header was parsed when the coded buffer was queued and its references + * still point into that buffer, which stays mapped until the job completes. + * + * Return: 0 with the hardware started and the watchdog armed, or a negative + * errno for a frame that cannot be decoded. + */ +static int rkjpegd_vdpu720_run(struct rkjpegd_ctx *ctx) +{ + struct rkjpegd_dev *jpegd = ctx->dev; + struct vb2_v4l2_buffer *src_buf, *dst_buf; + struct rkjpegd_src_buf *src; + const struct v4l2_jpeg_header *hdr; + dma_addr_t src_dma, dst_dma; + u32 data_offset, payload; + u32 hw_strm_off, strm_off, strm_start_byte, strm_len_blks; + int ret; + + src_buf = v4l2_m2m_next_src_buf(ctx->fh.m2m_ctx); + dst_buf = v4l2_m2m_next_dst_buf(ctx->fh.m2m_ctx); + src = vb2_to_rkjpegd_src_buf(&src_buf->vb2_buf); + + if (!src->parsed) + return -EINVAL; + + hdr = &src->header; + src_dma = vb2_dma_contig_plane_dma_addr(&src_buf->vb2_buf, 0); + dst_dma = vb2_dma_contig_plane_dma_addr(&dst_buf->vb2_buf, 0); + data_offset = src_buf->vb2_buf.planes[0].data_offset; + payload = vb2_get_plane_payload(&src_buf->vb2_buf, 0); + + /* + * Rebuild the Q/H-table side buffer every frame. table_base is a + * cached (dma_alloc_noncoherent) buffer and write_qtbl()/write_htbl() + * stage the source tables through cached stack buffers, so the build + * stays on cached memory; dma_sync_single_for_device() then flushes it + * to DRAM before the hardware reads it. + */ + memset(ctx->table_base.cpu, 0, ctx->table_base.size); + + ret = vdpu720_write_qtbl(ctx, hdr); + if (ret) + return ret; + + ret = vdpu720_write_htbl(ctx, hdr); + if (ret) + return ret; + + dma_sync_single_for_device(jpegd->dev, ctx->table_base.dma, + ctx->table_base.size, DMA_TO_DEVICE); + + /* + * The stream register must be 16-byte aligned. Round down to the + * nearest 16-byte boundary and record the sub-block start byte. + * + * Both are taken from the start of the plane rather than from the + * payload. data_offset is set by userspace in VIDIOC_QBUF and + * videobuf2 only rejects it when it is not smaller than bytesused, + * so it carries arbitrary low bits. Splitting an address that + * already includes it would leave STRM_BASE unaligned by those bits + * with no way to encode them, and the hardware would start reading + * from the wrong offset. + */ + strm_off = data_offset + hdr->ecs_offset; + hw_strm_off = strm_off & ~0xfU; + strm_start_byte = strm_off & 0xfU; + strm_len_blks = (ALIGN(payload - hw_strm_off, 16) - 1) >> 4; + + ret = vdpu720_fill_regs(ctx, hdr, ctx->table_base.dma, + src_dma + hw_strm_off, strm_start_byte, + strm_len_blks, dst_dma); + if (ret) + return ret; + + if (hdr->frame.num_components == 1) { + ret = vdpu720_fill_chroma(ctx, dst_buf); + if (ret) + return ret; + } + + rkjpegd_arm_watchdog(jpegd); + + /* + * Enable the buffer empty condition along with the timeout. A frame + * whose entropy data ends before the last macroblock row runs the + * decoder off the end of the stream, and with the condition masked it + * has nothing to report and waits. VDPU720_ERR_MASK already treats + * the resulting status as an error, so the handler is prepared for a + * status the hardware was never told to raise. + */ + rkjpegd_write(jpegd, + VDPU720_DEC_E | VDPU720_TIMEOUT_E | VDPU720_BUF_EMPTY_E, + VDPU720_REG_INT); + + return 0; +} + +/** + * vdpu720_soft_reset() - ask the block to reset itself + * @jpegd: device to reset + * + * Trigger the in-block soft reset and wait for it to report ready. Sleeps + * while polling, so it must not be called from atomic context. + * + * Return: 0 once the block reports the reset complete, -ETIMEDOUT if it does + * not do so within 10 ms. + */ +static int vdpu720_soft_reset(struct rkjpegd_dev *jpegd) +{ + u32 status; + int ret; + + /* + * If the decoder is idle (DEC_E=0), set FORCE_SOFTRESET_VALID + * before triggering the soft reset, per downstream BSP behaviour. + */ + status = rkjpegd_read(jpegd, VDPU720_REG_INT); + if (!(status & VDPU720_DEC_E)) + rkjpegd_write(jpegd, VDPU720_FORCE_SOFTRST, VDPU720_REG_SYS); + + rkjpegd_write(jpegd, status | VDPU720_SOFT_RST_EN, VDPU720_REG_INT); + + ret = readl_relaxed_poll_timeout(jpegd->regs + VDPU720_REG_INT, status, + status & VDPU720_SOFT_RST_RDY, + 5, 10000); + if (ret) + dev_warn(jpegd->dev, "soft reset timed out\n"); + + return ret; +} + +/** + * vdpu720_hard_reset() - pulse the block's reset lines + * @jpegd: device to reset + * + * reset_control_reset() is not usable here. The lines come from a Rockchip + * CRU, and rockchip_softrst_ops in drivers/clk/rockchip/softrst.c implements + * only .assert and .deassert, so reset_control_reset() returns -ENOTSUPP + * without touching the hardware. Drive the pulse by hand instead. + * + * Only reached from the watchdog, which runs from a workqueue, so sleeping + * between the two halves is fine. + * + * Return: 0 on success, negative errno if a reset control could not be + * asserted or deasserted. + */ +static int vdpu720_hard_reset(struct rkjpegd_dev *jpegd) +{ + int ret; + + ret = reset_control_assert(jpegd->resets); + if (ret) + return ret; + + usleep_range(10, 20); + + return reset_control_deassert(jpegd->resets); +} + +/** + * rkjpegd_vdpu720_reset() - error recovery reset called by the watchdog + * @ctx: context whose job timed out + * + * Try a soft reset first; fall back to a full hardware reset if the soft + * reset does not complete. + */ +static void rkjpegd_vdpu720_reset(struct rkjpegd_ctx *ctx) +{ + struct rkjpegd_dev *jpegd = ctx->dev; + int ret; + + ret = vdpu720_soft_reset(jpegd); + if (ret) { + dev_warn(jpegd->dev, "falling back to hard reset\n"); + + ret = vdpu720_hard_reset(jpegd); + if (ret) + dev_err(jpegd->dev, "hard reset failed: %d\n", ret); + } + + rkjpegd_write(jpegd, 0, VDPU720_REG_INT); +} + +static irqreturn_t rkjpegd_vdpu720_irq(int irq, void *dev_id) +{ + struct rkjpegd_dev *jpegd = dev_id; + enum vb2_buffer_state state; + u32 status, clr_mask; + + status = rkjpegd_read(jpegd, VDPU720_REG_INT); + + /* + * Two-phase IRQ clear. Write back a masked subset of status bits + * before checking IRQ_RAW, which some revisions require. + */ + clr_mask = (~(VDPU720_IRQ_CLR_COND & status)) & + (VDPU720_IRQ_CLR_KEEP & status); + rkjpegd_write(jpegd, clr_mask, VDPU720_REG_INT); + + if (!(status & VDPU720_IRQ_RAW)) + return IRQ_NONE; + + /* Fully clear IRQ */ + rkjpegd_write(jpegd, 0, VDPU720_REG_INT); + + state = (status & VDPU720_ERR_MASK) ? + VB2_BUF_STATE_ERROR : VB2_BUF_STATE_DONE; + + if (status & VDPU720_DEC_ERR) { + u32 mcu_pos = rkjpegd_read(jpegd, VDPU720_REG_DBG_MCU_POS); + u32 err_info = rkjpegd_read(jpegd, VDPU720_REG_DBG_ERROR); + + dev_warn_ratelimited(jpegd->dev, + "decode error: MCU pos=(%u,%u) flags=0x%04x [%s%s%s%s%s%s%s%s%s%s] first_idx=%u\n", + (u32)FIELD_GET(VDPU720_DBG_MCU_POS_X, mcu_pos), + (u32)FIELD_GET(VDPU720_DBG_MCU_POS_Y, mcu_pos), + (u32)FIELD_GET(VDPU720_DERR_FLAGS, err_info), + (err_info & VDPU720_DERR_DRI_SEQ) ? "dri_seq " : "", + (err_info & VDPU720_DERR_STREAM_FFFF) ? "ffff " : "", + (err_info & VDPU720_DERR_OTHER_MARK) ? "bad_mark " : "", + (err_info & VDPU720_DERR_MCU_CNT_L) ? "dri_early " : "", + (err_info & VDPU720_DERR_MCU_CNT_M) ? "dri_late " : "", + (err_info & VDPU720_DERR_EOI_NO_END) ? "eoi_early " : "", + (err_info & VDPU720_DERR_END_NO_EOI) ? "no_eoi " : "", + (err_info & VDPU720_DERR_OVERFLOW) ? "overflow " : "", + (err_info & VDPU720_DERR_HUFF_EMPTY) ? "huff_empty " : "", + (err_info & (VDPU720_DERR_STREAM_R0 | + VDPU720_DERR_STREAM_R1)) ? "stream_mark " : "", + (u32)FIELD_GET(VDPU720_DERR_FIRST_IDX, err_info)); + + /* Clear the sticky flags so they do not bleed into the next frame */ + rkjpegd_write(jpegd, err_info, VDPU720_REG_DBG_ERROR); + } + + rkjpegd_irq_done(jpegd, state); + + return IRQ_HANDLED; +} + +static void rkjpegd_watchdog(struct work_struct *work) +{ + struct rkjpegd_dev *jpegd = container_of(to_delayed_work(work), + struct rkjpegd_dev, + watchdog_work); + struct rkjpegd_ctx *ctx = v4l2_m2m_get_curr_priv(jpegd->m2m_dev); + + if (!ctx) + return; + + dev_err(jpegd->dev, "frame processing timed out\n"); + rkjpegd_vdpu720_reset(ctx); + rkjpegd_job_finish(ctx, VB2_BUF_STATE_ERROR); +} + +/** + * rkjpegd_source_change() - report a resolution change to userspace + * @ctx: context the coded buffer belongs to + * @src_buf: coded buffer whose header the new resolution is taken from + * + * Renegotiates the capture format from @src_buf and reports it, unless it + * already describes what was negotiated. Sets @rkjpegd_ctx.source_change, + * which stops rkjpegd_job_ready() from letting any further job run until the + * capture queue has been set up again. + * + * Called both when a coded buffer is queued and when one reaches the head of + * the queue, see the comments at those two call sites. + */ +static void rkjpegd_source_change(struct rkjpegd_ctx *ctx, + struct rkjpegd_src_buf *src_buf) +{ + u32 width, height, buf_width, buf_height; + + if (src_buf->parsed) { + width = src_buf->header.frame.width; + height = src_buf->header.frame.height; + } else { + /* + * A frame the parser refused carries no dimensions of its own, + * and the job will hand it back with an error. Report the + * change anyway, from the coded format userspace configured. + * V4L2_FMT_FLAG_DYN_RESOLUTION tells an application to wait + * for this event before it sets the capture queue up, so one + * that never arrives leaves it waiting for a frame that cannot + * come instead of seeing the error and giving up. + */ + width = ctx->src_fmt.width; + height = ctx->src_fmt.height; + } + buf_width = ALIGN(width, RKJPEGD_RAW_STEP); + buf_height = ALIGN(height, RKJPEGD_RAW_STEP); + + if (!ctx->initial_source_change && + ctx->dst_fmt.width == buf_width && + ctx->dst_fmt.height == buf_height && + ctx->crop.width == width && ctx->crop.height == height) + return; + + dev_dbg(ctx->dev->dev, "source change to %ux%u\n", width, height); + + rkjpegd_fill_raw_fmt(&ctx->dst_fmt, buf_width, buf_height); + ctx->crop.left = 0; + ctx->crop.top = 0; + ctx->crop.width = width; + ctx->crop.height = height; + + v4l2_event_queue_fh(&ctx->fh, &rkjpegd_src_change_event); + ctx->source_change = true; + ctx->initial_source_change = false; + + if (vb2_is_streaming(v4l2_m2m_get_dst_vq(ctx->fh.m2m_ctx))) + rkjpegd_set_last_buffer(ctx); +} + +static void rkjpegd_device_run(void *priv) +{ + struct rkjpegd_ctx *ctx = priv; + struct rkjpegd_dev *jpegd = ctx->dev; + struct vb2_v4l2_buffer *src, *dst; + int ret; + + src = v4l2_m2m_next_src_buf(ctx->fh.m2m_ctx); + dst = v4l2_m2m_next_dst_buf(ctx->fh.m2m_ctx); + if (WARN_ON(!src) || WARN_ON(!dst)) + return; + + /* + * rkjpegd_buf_queue() only looks for a resolution change while the + * coded queue is empty, so a change carried by a buffer queued behind + * others is not seen there. Look again now that this buffer has + * reached the head: everything queued before it has been decoded and + * handed back, which is the point the specification wants the change + * reported at. + * + * Doing it here rather than earlier is also what keeps the buffers + * queued behind this one correct. They are still coded frames of the + * new resolution, and they stay queued while the capture queue is set + * up again, so they are decoded into buffers that fit them. + */ + rkjpegd_source_change(ctx, vb2_to_rkjpegd_src_buf(&src->vb2_buf)); + if (ctx->source_change) { + /* + * rkjpegd_set_last_buffer() has given dst back to userspace + * with V4L2_BUF_FLAG_LAST set, so it must not be touched here. + * Finish the job without consuming src either: it is decoded + * once the capture queue has been set up for it. + */ + v4l2_m2m_job_finish(jpegd->m2m_dev, ctx->fh.m2m_ctx); + return; + } + + ret = pm_runtime_resume_and_get(jpegd->dev); + if (ret < 0) + goto err_finish; + + ret = clk_bulk_enable(RKJPEGD_NUM_CLOCKS, jpegd->clocks); + if (ret) + goto err_pm_put; + + v4l2_m2m_buf_copy_metadata(src, dst); + + ret = rkjpegd_vdpu720_run(ctx); + if (ret) + goto err_clk_disable; + + return; + +err_clk_disable: + clk_bulk_disable(RKJPEGD_NUM_CLOCKS, jpegd->clocks); +err_pm_put: + pm_runtime_put_autosuspend(jpegd->dev); +err_finish: + rkjpegd_job_finish_no_pm(ctx, VB2_BUF_STATE_ERROR); +} + +static int rkjpegd_job_ready(void *priv) +{ + struct rkjpegd_ctx *ctx = priv; + + return ctx->source_change ? 0 : 1; +} + +static const struct v4l2_m2m_ops rkjpegd_m2m_ops = { + .device_run = rkjpegd_device_run, + .job_ready = rkjpegd_job_ready, +}; + +/* + * Bitstream inspection + * + * The header is parsed when the buffer is queued rather than when the job + * runs: the resolution it carries is what a source change reports, and the + * references it hands out point into the payload, which stays mapped until + * the buffer is given back. + */ + +static bool rkjpegd_has_eoi(const void *data, u32 len) +{ + u8 tail[64]; + u32 tail_len, i; + + /* + * v4l2_jpeg_parse_header() stops at the start of scan marker and never + * looks at the entropy coded data behind it, so a frame truncated + * because it did not fit the buffer parses without an error. The + * hardware would decode as many macroblocks as it finds and hand out a + * half filled frame indistinguishable from a good one. + * + * An end of image marker cannot appear inside the entropy coded data, + * where 0xff bytes are stuffed, so finding one means the frame is + * complete. Only the last few bytes are searched, which covers a + * payload padded up to a 64 byte boundary and any short trailer. + * Looking further back is not worth it: the buffer is uncached, so a + * walk over the whole payload costs one bus transaction per byte, and + * a payload ending far behind its marker is one whose bytesused was + * never set, in which case videobuf2 substitutes the full plane length + * and a recycled buffer holds the previous frame's bytes back there + * anyway. The tail is copied out in one memcpy() for the same reason. + */ + tail_len = min_t(u32, len, sizeof(tail)); + memcpy(tail, data + len - tail_len, tail_len); + + for (i = 0; i + 1 < tail_len; i++) + if (tail[i] == 0xff && tail[i + 1] == 0xd9) + return true; + + return false; +} + +static bool rkjpegd_header_supported(struct rkjpegd_dev *jpegd, + const struct v4l2_jpeg_header *header, + u32 len) +{ + /* + * v4l2_jpeg_parse_header() accepts twelve bit samples for SOF1, but + * the register programming always asks for eight bit ones. + */ + if (header->frame.precision != 8) { + dev_err_ratelimited(jpegd->dev, + "unsupported JPEG sample precision %u\n", + header->frame.precision); + return false; + } + + /* + * Only a single component frame and the three component layouts are + * decodable. The hardware mode is derived from the luma sampling + * factors alone, so a two component frame would come back as one of + * the three component modes and the hardware would go looking for a + * chroma plane that is not in the bitstream. + */ + if (header->frame.num_components != 1 && + header->frame.num_components != 3) { + dev_err_ratelimited(jpegd->dev, + "unsupported JPEG component count %u\n", + header->frame.num_components); + return false; + } + + /* + * The decoder runs a whole frame in one go, so the scan has to cover + * every component. A non interleaved frame, whose first scan carries + * a single component, would leave the rest of the table side buffer + * zeroed. + */ + if (header->scan->num_components != header->frame.num_components) { + dev_err_ratelimited(jpegd->dev, + "JPEG scan covers %u of %u components, non interleaved scans are not supported\n", + header->scan->num_components, + header->frame.num_components); + return false; + } + + if (header->ecs_offset >= len) { + dev_err_ratelimited(jpegd->dev, + "JPEG entropy coded data starts beyond the payload\n"); + return false; + } + + return true; +} + +static void rkjpegd_parse_src_buf(struct rkjpegd_ctx *ctx, + struct vb2_buffer *vb) +{ + struct rkjpegd_src_buf *src_buf = vb2_to_rkjpegd_src_buf(vb); + struct rkjpegd_dev *jpegd = ctx->dev; + u32 data_offset = vb->planes[0].data_offset; + u32 len = vb2_get_plane_payload(vb, 0); + void *data = vb2_plane_vaddr(vb, 0); + int ret; + + memset(&src_buf->header, 0, sizeof(src_buf->header)); + memset(&src_buf->scan, 0, sizeof(src_buf->scan)); + memset(src_buf->quantization_tables, 0, + sizeof(src_buf->quantization_tables)); + memset(src_buf->huffman_tables, 0, sizeof(src_buf->huffman_tables)); + src_buf->header.scan = &src_buf->scan; + src_buf->header.quantization_tables = src_buf->quantization_tables; + src_buf->header.huffman_tables = src_buf->huffman_tables; + src_buf->parsed = false; + + if (!data) { + dev_err_ratelimited(jpegd->dev, + "JPEG buffer has no kernel mapping\n"); + return; + } + + if (len <= data_offset || len - data_offset < 4) { + dev_err_ratelimited(jpegd->dev, + "JPEG payload of %u bytes is too short\n", + len); + return; + } + + data += data_offset; + len -= data_offset; + + ret = v4l2_jpeg_parse_header(data, len, &src_buf->header); + if (ret < 0) { + dev_warn_ratelimited(jpegd->dev, + "failed to parse JPEG header: %d (len=%u first_bytes=%*ph)\n", + ret, len, min_t(int, len, 8), data); + return; + } + + if (!rkjpegd_header_supported(jpegd, &src_buf->header, len)) + return; + + if (!rkjpegd_has_eoi(data, len)) { + dev_err_ratelimited(jpegd->dev, + "truncated JPEG, no end of image marker at the end of the %u byte payload (buffer too small?)\n", + len); + return; + } + + src_buf->parsed = true; +} + +/* + * videobuf2 operations + */ + +static int rkjpegd_queue_setup(struct vb2_queue *vq, unsigned int *num_buffers, + unsigned int *num_planes, unsigned int sizes[], + struct device *alloc_devs[]) +{ + struct rkjpegd_ctx *ctx = vb2_get_drv_priv(vq); + struct v4l2_pix_format_mplane *pix_mp; + + pix_mp = V4L2_TYPE_IS_OUTPUT(vq->type) ? &ctx->src_fmt : &ctx->dst_fmt; + + if (*num_planes) { + if (*num_planes != 1) + return -EINVAL; + if (sizes[0] < pix_mp->plane_fmt[0].sizeimage) + return -EINVAL; + return 0; + } + + *num_planes = 1; + sizes[0] = pix_mp->plane_fmt[0].sizeimage; + + /* + * The first frame to arrive after the coded queue was set up has to + * report its resolution even when it happens to match what userspace + * guessed, otherwise an application waiting for the event before it + * allocates capture buffers never gets one. + */ + if (V4L2_TYPE_IS_OUTPUT(vq->type)) + ctx->initial_source_change = true; + + return 0; +} + +static int rkjpegd_buf_out_validate(struct vb2_buffer *vb) +{ + struct vb2_v4l2_buffer *vbuf = to_vb2_v4l2_buffer(vb); + + vbuf->field = V4L2_FIELD_NONE; + + return 0; +} + +static int rkjpegd_buf_prepare(struct vb2_buffer *vb) +{ + struct vb2_queue *vq = vb->vb2_queue; + struct rkjpegd_ctx *ctx = vb2_get_drv_priv(vq); + struct v4l2_pix_format_mplane *pix_mp; + + if (V4L2_TYPE_IS_OUTPUT(vq->type)) { + pix_mp = &ctx->src_fmt; + + if (vb2_plane_size(vb, 0) < pix_mp->plane_fmt[0].sizeimage) + return -EINVAL; + + return 0; + } + + pix_mp = &ctx->dst_fmt; + + /* + * While a source change is pending the capture buffers still have the + * size of the previous format and are about to be reallocated, and no + * job runs until they have been. Do not reject them here - but do not + * describe them with the new format either, which is larger than they + * are whenever the resolution grew, and would have + * vb2_set_plane_payload() warn and clamp. They carry no picture, so + * leave the payload at zero and let them be given back empty. + */ + if (ctx->source_change) + return 0; + + if (vb2_plane_size(vb, 0) < pix_mp->plane_fmt[0].sizeimage) + return -EINVAL; + + vb2_set_plane_payload(vb, 0, pix_mp->plane_fmt[0].sizeimage); + + return 0; +} + +static void rkjpegd_buf_queue(struct vb2_buffer *vb) +{ + struct vb2_v4l2_buffer *vbuf = to_vb2_v4l2_buffer(vb); + struct rkjpegd_ctx *ctx = vb2_get_drv_priv(vb->vb2_queue); + + if (V4L2_TYPE_IS_CAPTURE(vb->vb2_queue->type)) { + if (vb2_is_streaming(vb->vb2_queue) && + v4l2_m2m_dst_buf_is_last(ctx->fh.m2m_ctx)) { + vbuf->field = V4L2_FIELD_NONE; + vbuf->sequence = ctx->sequence_cap++; + v4l2_m2m_last_buffer_done(ctx->fh.m2m_ctx, vbuf); + v4l2_event_queue_fh(&ctx->fh, &rkjpegd_eos_event); + return; + } + + v4l2_m2m_buf_queue(ctx->fh.m2m_ctx, vbuf); + return; + } + + rkjpegd_parse_src_buf(ctx, vb); + + /* + * Only look at a resolution change while nothing else is pending. A + * buffer queued behind frames that have not been decoded yet would + * otherwise renegotiate the capture queue out from under them; + * rkjpegd_device_run() picks that case up once the buffer reaches the + * head of the queue. + * + * This early report is still needed for the first buffer of a stream. + * The mem2mem core runs no job until both queues are streaming, so + * rkjpegd_device_run() cannot report the resolution an application has + * to know before it can set the capture queue up and stream it on. + */ + if (!v4l2_m2m_num_src_bufs_ready(ctx->fh.m2m_ctx)) + rkjpegd_source_change(ctx, vb2_to_rkjpegd_src_buf(vb)); + + v4l2_m2m_buf_queue(ctx->fh.m2m_ctx, vbuf); +} + +static int rkjpegd_start_streaming(struct vb2_queue *vq, unsigned int count) +{ + struct rkjpegd_ctx *ctx = vb2_get_drv_priv(vq); + + v4l2_m2m_update_start_streaming_state(ctx->fh.m2m_ctx, vq); + + if (V4L2_TYPE_IS_OUTPUT(vq->type)) { + ctx->sequence_out = 0; + } else { + ctx->sequence_cap = 0; + ctx->source_change = false; + ctx->initial_source_change = false; + } + + return 0; +} + +static void rkjpegd_stop_streaming(struct vb2_queue *vq) +{ + struct rkjpegd_ctx *ctx = vb2_get_drv_priv(vq); + struct vb2_v4l2_buffer *vbuf; + + for (;;) { + if (V4L2_TYPE_IS_OUTPUT(vq->type)) + vbuf = v4l2_m2m_src_buf_remove(ctx->fh.m2m_ctx); + else + vbuf = v4l2_m2m_dst_buf_remove(ctx->fh.m2m_ctx); + if (!vbuf) + break; + v4l2_m2m_buf_done(vbuf, VB2_BUF_STATE_ERROR); + } + + v4l2_m2m_update_stop_streaming_state(ctx->fh.m2m_ctx, vq); + + /* + * A drain that was requested before the source change was reported + * has not run yet; keep it pending across the capture queue restart. + */ + if (V4L2_TYPE_IS_CAPTURE(vq->type) && ctx->source_change && + ctx->fh.m2m_ctx->last_src_buf) + ctx->fh.m2m_ctx->is_draining = true; + + if (V4L2_TYPE_IS_OUTPUT(vq->type) && + v4l2_m2m_has_stopped(ctx->fh.m2m_ctx)) + v4l2_event_queue_fh(&ctx->fh, &rkjpegd_eos_event); +} + +static const struct vb2_ops rkjpegd_queue_ops = { + .queue_setup = rkjpegd_queue_setup, + .buf_out_validate = rkjpegd_buf_out_validate, + .buf_prepare = rkjpegd_buf_prepare, + .buf_queue = rkjpegd_buf_queue, + .start_streaming = rkjpegd_start_streaming, + .stop_streaming = rkjpegd_stop_streaming, +}; + +static int rkjpegd_queue_init(void *priv, struct vb2_queue *src_vq, + struct vb2_queue *dst_vq) +{ + struct rkjpegd_ctx *ctx = priv; + struct rkjpegd_dev *jpegd = ctx->dev; + int ret; + + src_vq->type = V4L2_BUF_TYPE_VIDEO_OUTPUT_MPLANE; + src_vq->io_modes = VB2_MMAP | VB2_DMABUF; + src_vq->drv_priv = ctx; + src_vq->ops = &rkjpegd_queue_ops; + src_vq->mem_ops = &vb2_dma_contig_memops; + src_vq->buf_struct_size = sizeof(struct rkjpegd_src_buf); + src_vq->timestamp_flags = V4L2_BUF_FLAG_TIMESTAMP_COPY; + src_vq->lock = &jpegd->vdev_lock; + src_vq->dev = jpegd->v4l2_dev.dev; + + /* + * Driver does mostly sequential access, so sacrifice TLB efficiency + * for faster allocation. Both queues keep their kernel mapping: the + * quantisation and Huffman tables are read out of the coded buffer, + * and a grayscale frame leaves its chroma plane to the driver. + */ + src_vq->dma_attrs = DMA_ATTR_ALLOC_SINGLE_PAGES; + + ret = vb2_queue_init(src_vq); + if (ret) + return ret; + + dst_vq->type = V4L2_BUF_TYPE_VIDEO_CAPTURE_MPLANE; + dst_vq->io_modes = VB2_MMAP | VB2_DMABUF; + dst_vq->drv_priv = ctx; + dst_vq->ops = &rkjpegd_queue_ops; + dst_vq->mem_ops = &vb2_dma_contig_memops; + dst_vq->buf_struct_size = sizeof(struct v4l2_m2m_buffer); + dst_vq->timestamp_flags = V4L2_BUF_FLAG_TIMESTAMP_COPY; + dst_vq->lock = &jpegd->vdev_lock; + dst_vq->dev = jpegd->v4l2_dev.dev; + dst_vq->dma_attrs = DMA_ATTR_ALLOC_SINGLE_PAGES; + + return vb2_queue_init(dst_vq); +} + +/* + * File operations + */ + +static int rkjpegd_open(struct file *filp) +{ + struct rkjpegd_dev *jpegd = video_drvdata(filp); + struct rkjpegd_ctx *ctx; + int ret; + + ctx = kzalloc_obj(*ctx); + if (!ctx) + return -ENOMEM; + + ctx->dev = jpegd; + rkjpegd_reset_fmts(ctx); + v4l2_fh_init(&ctx->fh, video_devdata(filp)); + + ctx->fh.m2m_ctx = v4l2_m2m_ctx_init(jpegd->m2m_dev, ctx, + rkjpegd_queue_init); + if (IS_ERR(ctx->fh.m2m_ctx)) { + ret = PTR_ERR(ctx->fh.m2m_ctx); + goto err_free_ctx; + } + + ret = rkjpegd_vdpu720_init(ctx); + if (ret) + goto err_cleanup_m2m_ctx; + + v4l2_fh_add(&ctx->fh, filp); + + return 0; + +err_cleanup_m2m_ctx: + v4l2_m2m_ctx_release(ctx->fh.m2m_ctx); +err_free_ctx: + v4l2_fh_exit(&ctx->fh); + kfree(ctx); + + return ret; +} + +static int rkjpegd_release(struct file *filp) +{ + struct rkjpegd_ctx *ctx = file_to_rkjpegd_ctx(filp); + + v4l2_fh_del(&ctx->fh, filp); + v4l2_m2m_ctx_release(ctx->fh.m2m_ctx); + rkjpegd_vdpu720_exit(ctx); + v4l2_fh_exit(&ctx->fh); + kfree(ctx); + + return 0; +} + +static const struct v4l2_file_operations rkjpegd_fops = { + .owner = THIS_MODULE, + .open = rkjpegd_open, + .release = rkjpegd_release, + .poll = v4l2_m2m_fop_poll, + .unlocked_ioctl = video_ioctl2, + .mmap = v4l2_m2m_fop_mmap, +}; + +/* + * Device setup + */ + +static int rkjpegd_v4l2_init(struct rkjpegd_dev *jpegd) +{ + int ret; + + ret = v4l2_device_register(jpegd->dev, &jpegd->v4l2_dev); + if (ret) { + dev_err(jpegd->dev, "failed to register V4L2 device\n"); + return ret; + } + + jpegd->m2m_dev = v4l2_m2m_init(&rkjpegd_m2m_ops); + if (IS_ERR(jpegd->m2m_dev)) { + v4l2_err(&jpegd->v4l2_dev, "failed to init mem2mem device\n"); + ret = PTR_ERR(jpegd->m2m_dev); + goto err_unregister_v4l2; + } + + jpegd->mdev.dev = jpegd->dev; + strscpy(jpegd->mdev.model, RKJPEGD_NAME, sizeof(jpegd->mdev.model)); + media_device_init(&jpegd->mdev); + jpegd->v4l2_dev.mdev = &jpegd->mdev; + + jpegd->vdev.lock = &jpegd->vdev_lock; + jpegd->vdev.v4l2_dev = &jpegd->v4l2_dev; + jpegd->vdev.fops = &rkjpegd_fops; + jpegd->vdev.release = video_device_release_empty; + jpegd->vdev.vfl_dir = VFL_DIR_M2M; + jpegd->vdev.device_caps = V4L2_CAP_STREAMING | + V4L2_CAP_VIDEO_M2M_MPLANE; + jpegd->vdev.ioctl_ops = &rkjpegd_ioctl_ops; + video_set_drvdata(&jpegd->vdev, jpegd); + strscpy(jpegd->vdev.name, RKJPEGD_NAME, sizeof(jpegd->vdev.name)); + + ret = video_register_device(&jpegd->vdev, VFL_TYPE_VIDEO, -1); + if (ret) { + v4l2_err(&jpegd->v4l2_dev, "failed to register video device\n"); + goto err_cleanup_mc; + } + + ret = v4l2_m2m_register_media_controller(jpegd->m2m_dev, &jpegd->vdev, + MEDIA_ENT_F_PROC_VIDEO_DECODER); + if (ret) { + v4l2_err(&jpegd->v4l2_dev, + "failed to init V4L2 M2M media controller\n"); + goto err_unregister_vdev; + } + + ret = media_device_register(&jpegd->mdev); + if (ret) { + v4l2_err(&jpegd->v4l2_dev, "failed to register media device\n"); + goto err_unregister_mc; + } + + return 0; + +err_unregister_mc: + v4l2_m2m_unregister_media_controller(jpegd->m2m_dev); +err_unregister_vdev: + video_unregister_device(&jpegd->vdev); +err_cleanup_mc: + media_device_cleanup(&jpegd->mdev); + v4l2_m2m_release(jpegd->m2m_dev); +err_unregister_v4l2: + v4l2_device_unregister(&jpegd->v4l2_dev); + + return ret; +} + +static void rkjpegd_v4l2_cleanup(struct rkjpegd_dev *jpegd) +{ + media_device_unregister(&jpegd->mdev); + v4l2_m2m_unregister_media_controller(jpegd->m2m_dev); + video_unregister_device(&jpegd->vdev); + media_device_cleanup(&jpegd->mdev); + v4l2_m2m_release(jpegd->m2m_dev); + v4l2_device_unregister(&jpegd->v4l2_dev); +} + +static int rkjpegd_probe(struct platform_device *pdev) +{ + struct rkjpegd_dev *jpegd; + unsigned int i; + int irq, ret; + + jpegd = devm_kzalloc(&pdev->dev, sizeof(*jpegd), GFP_KERNEL); + if (!jpegd) + return -ENOMEM; + + jpegd->dev = &pdev->dev; + platform_set_drvdata(pdev, jpegd); + mutex_init(&jpegd->vdev_lock); + INIT_DELAYED_WORK(&jpegd->watchdog_work, rkjpegd_watchdog); + + for (i = 0; i < RKJPEGD_NUM_CLOCKS; i++) + jpegd->clocks[i].id = rkjpegd_clk_names[i]; + + ret = devm_clk_bulk_get(&pdev->dev, RKJPEGD_NUM_CLOCKS, jpegd->clocks); + if (ret) + return ret; + + jpegd->resets = devm_reset_control_array_get_exclusive(&pdev->dev); + if (IS_ERR(jpegd->resets)) + return dev_err_probe(&pdev->dev, PTR_ERR(jpegd->resets), + "failed to get resets\n"); + + jpegd->regs = devm_platform_ioremap_resource(pdev, 0); + if (IS_ERR(jpegd->regs)) + return PTR_ERR(jpegd->regs); + + ret = dma_set_mask_and_coherent(&pdev->dev, DMA_BIT_MASK(32)); + if (ret) + return dev_err_probe(&pdev->dev, ret, "failed to set DMA mask\n"); + + irq = platform_get_irq(pdev, 0); + if (irq < 0) + return irq; + + ret = devm_request_irq(&pdev->dev, irq, rkjpegd_vdpu720_irq, 0, + dev_name(&pdev->dev), jpegd); + if (ret) + return dev_err_probe(&pdev->dev, ret, "failed to request irq\n"); + + ret = clk_bulk_prepare(RKJPEGD_NUM_CLOCKS, jpegd->clocks); + if (ret) + return dev_err_probe(&pdev->dev, ret, + "failed to prepare clocks\n"); + + pm_runtime_set_autosuspend_delay(&pdev->dev, 100); + pm_runtime_use_autosuspend(&pdev->dev); + pm_runtime_enable(&pdev->dev); + + ret = reset_control_deassert(jpegd->resets); + if (ret) { + dev_err_probe(&pdev->dev, ret, "failed to deassert resets\n"); + goto err_disable_pm; + } + + ret = rkjpegd_v4l2_init(jpegd); + if (ret) + goto err_assert_resets; + + return 0; + +err_assert_resets: + reset_control_assert(jpegd->resets); +err_disable_pm: + pm_runtime_dont_use_autosuspend(&pdev->dev); + pm_runtime_disable(&pdev->dev); + clk_bulk_unprepare(RKJPEGD_NUM_CLOCKS, jpegd->clocks); + + return ret; +} + +static void rkjpegd_remove(struct platform_device *pdev) +{ + struct rkjpegd_dev *jpegd = platform_get_drvdata(pdev); + + cancel_delayed_work_sync(&jpegd->watchdog_work); + rkjpegd_v4l2_cleanup(jpegd); + reset_control_assert(jpegd->resets); + pm_runtime_dont_use_autosuspend(&pdev->dev); + pm_runtime_disable(&pdev->dev); + clk_bulk_unprepare(RKJPEGD_NUM_CLOCKS, jpegd->clocks); +} + +/* + * The clocks are prepared once and only enabled and disabled around a job, + * which is what lets rkjpegd_job_finish() run from the interrupt handler. + * There is nothing else to do around a runtime transition: the power domain + * is handled by genpd on the device's behalf. + */ +static const struct dev_pm_ops rkjpegd_pm_ops = { + SYSTEM_SLEEP_PM_OPS(pm_runtime_force_suspend, pm_runtime_force_resume) +}; + +static const struct of_device_id of_rkjpegd_match[] = { + { .compatible = "rockchip,rk3568-jpegd" }, + { .compatible = "rockchip,rk3588-jpegd" }, + { /* sentinel */ } +}; +MODULE_DEVICE_TABLE(of, of_rkjpegd_match); + +static struct platform_driver rkjpegd_driver = { + .probe = rkjpegd_probe, + .remove = rkjpegd_remove, + .driver = { + .name = RKJPEGD_NAME, + .of_match_table = of_rkjpegd_match, + .pm = pm_sleep_ptr(&rkjpegd_pm_ops), + }, +}; +module_platform_driver(rkjpegd_driver); + +MODULE_DESCRIPTION("Rockchip JPEG decoder driver"); +MODULE_AUTHOR("Lucas Sinn <[email protected]>"); +MODULE_LICENSE("GPL"); -- 2.47.3