๐Ÿงฎ Zenith TX26 โ€” 16-bit CPU

A 16-bit processor written from scratch in Verilog โ€” no cores, no generators, no IP blocks. Nine modules, sixteen instructions, verified on a Tang Nano 9K FPGA and submitted to TinyTapeout for fabrication.

Taped out via TinyTapeout Physical silicon expected back Q4 2026. The design below is what's on the die.
16-bit datapath Harvard architecture Single-cycle 8 registers (R0โ€“R7) 16 instructions 9 Verilog modules Icarus Verilog + GTKWave Tang Nano 9K

๐Ÿ”€ Datapath

Instruction and data memory are kept separate, so a fetch and a data access never contend for the same port โ€” that's what makes single-cycle execution possible without stalling. Every instruction completes in one clock edge: the program counter addresses the instruction ROM, the decoder splits the 16-bit word into its fields, the register file reads two operands combinationally, the ALU computes, and the result is written back before the next rising edge.

R0 is hardwired to zero. That one decision removes the need for separate move and clear instructions โ€” ADD R1, R2, R0 is a register copy, and comparing against R0 is a test for zero. It's the same trick MIPS and RISC-V use, and it's why 16 opcodes are enough.

Zenith TX26 datapath overview The program counter addresses the instruction memory. The fetched 16-bit instruction goes to the instruction decoder, which splits it into opcode and register fields. The decoder drives the control unit and the sign extender. The register file reads two operands and feeds the ALU. The ALU result is written back to the register file, and also serves as the address for data memory on load and store instructions. Branch and jump targets feed back into the program counter. instruction [15:0] ra, rb opcode [15:12] ALU op / write enable branch / jump target imm [5:0] imm [15:0] addr / data write-back โ†’ rd clk / reset Program Counter Instruction ROM Harvard: separate bus Decoder splits R / I / J fields Control Unit FSM Sign Extender 6-bit โ†’ 16-bit Register File 8 ร— 16-bit R0 hardwired to 0 ALU add, sub, and, or, xor, shift, compare Data RAM LD / ST only Data path Control signals Program flow

๐Ÿ“ Instruction Encoding

Every instruction is exactly 16 bits wide, in one of three layouts. The opcode always occupies the top four bits, so the decoder can identify the format before it knows anything else about the instruction.

R-type โ€” register operations
15:12opcode
11:9rd
8:6ra
5:3rb
2:0reserved
I-type โ€” immediate operations
15:12opcode
11:9rd
8:6ra
5:0immediate, sign-extended
J-type โ€” jumps and branches
15:12opcode
11:0immediate

๐Ÿ“‹ Instruction Set

OpcodeTypeInstructionOperation
0000RADD rd, ra, rbrd = ra + rb
0001RSUB rd, ra, rbrd = ra โˆ’ rb
0010RAND rd, ra, rbrd = ra & rb
0011ROR rd, ra, rbrd = ra | rb
0100RXOR rd, ra, rbrd = ra ^ rb
0101RSHL rd, rard = ra << 1
0110RSHR rd, rard = ra >> 1
0111RSLT rd, ra, rbrd = (ra < rb) ? 1 : 0
1000ILDI rd, immrd = imm
1001IADDI rd, ra, immrd = ra + imm
1010ILD rd, ra, immrd = mem[ra + imm]
1011IST rd, ra, immmem[rd + imm] = ra
1100JJMP immpc = imm
1101JBEQ ra, rb, immif ra == rb: pc = pc + imm
1110JBLT ra, rb, immif ra < rb: pc = pc + imm
1111โ€”NOPdo nothing

โ–ถ๏ธ Running Code

Programs are written in assembly and turned into a memory image by assembler.py, a small two-pass assembler written for this instruction set. The program taped out with the chip is an endless Fibonacci loop โ€” it runs until the sequence overflows 16 bits and wraps, then keeps going.

LDI R1, 0        # R1 = 0   (current)
LDI R2, 1        # R2 = 1   (next)
NOP              # loop start
ADD R3, R1, R2   # temp = R1 + R2
NOP
ADD R1, R2, R0   # R1 = R2  (R0 is zero, so this is a copy)
NOP
ADD R2, R3, R0   # R2 = temp
JMP 2            # back to loop start

The NOPs are there because the register file writes on the clock edge โ€” a result isn't readable by the very next instruction. Spacing the dependent instructions apart is the simplest way to respect that on a design with no forwarding or interlocks. Building a CPU that can issue those instructions back to back means adding a hazard unit, and that's the next revision.

๐Ÿ”ฌ Verification

Each module was simulated in Icarus Verilog and inspected waveform by waveform in GTKWave before being wired into the top level. The load and store instructions were the last to work โ€” the store encoding uses rd as the address base rather than the destination, which is easy to get backwards and produces silently wrong memory writes rather than an obvious failure.

After simulation passed, the design was synthesized for a Tang Nano 9K FPGA. The board's six LEDs are driven by the low six bits of the program counter, so the loop is visible as a counting pattern โ€” real hardware executing the instruction set, not a testbench asserting that it should.

Icarus Verilog GTKWave Sipeed Tang Nano 9K Gowin toolchain TinyTapeout submission
View the Verilog on GitHub
โ„น

About Me

I'm Doruk Orak, a 12th grade student at Sankt Georg Austrian High School in Istanbul.

My interests span hardware design, artificial intelligence, programming, and physics. Connecting the world of intelligent systems with physical circuits.