A 16-bit processor written from scratch in Verilog โ no cores, no generators, no IP blocks. Nine modules, sixteen instructions, verified on a Tang Nano 9K FPGA and submitted to TinyTapeout for fabrication.
Instruction and data memory are kept separate, so a fetch and a data access never contend for the same port โ that's what makes single-cycle execution possible without stalling. Every instruction completes in one clock edge: the program counter addresses the instruction ROM, the decoder splits the 16-bit word into its fields, the register file reads two operands combinationally, the ALU computes, and the result is written back before the next rising edge.
R0 is hardwired to zero. That one decision removes the need for separate move and clear instructions โ ADD R1, R2, R0 is a register copy, and comparing against R0 is a test for zero. It's the same trick MIPS and RISC-V use, and it's why 16 opcodes are enough.
Every instruction is exactly 16 bits wide, in one of three layouts. The opcode always occupies the top four bits, so the decoder can identify the format before it knows anything else about the instruction.
| Opcode | Type | Instruction | Operation |
|---|---|---|---|
| 0000 | R | ADD rd, ra, rb | rd = ra + rb |
| 0001 | R | SUB rd, ra, rb | rd = ra โ rb |
| 0010 | R | AND rd, ra, rb | rd = ra & rb |
| 0011 | R | OR rd, ra, rb | rd = ra | rb |
| 0100 | R | XOR rd, ra, rb | rd = ra ^ rb |
| 0101 | R | SHL rd, ra | rd = ra << 1 |
| 0110 | R | SHR rd, ra | rd = ra >> 1 |
| 0111 | R | SLT rd, ra, rb | rd = (ra < rb) ? 1 : 0 |
| 1000 | I | LDI rd, imm | rd = imm |
| 1001 | I | ADDI rd, ra, imm | rd = ra + imm |
| 1010 | I | LD rd, ra, imm | rd = mem[ra + imm] |
| 1011 | I | ST rd, ra, imm | mem[rd + imm] = ra |
| 1100 | J | JMP imm | pc = imm |
| 1101 | J | BEQ ra, rb, imm | if ra == rb: pc = pc + imm |
| 1110 | J | BLT ra, rb, imm | if ra < rb: pc = pc + imm |
| 1111 | โ | NOP | do nothing |
Programs are written in assembly and turned into a memory image by assembler.py, a small two-pass assembler written for this instruction set. The program taped out with the chip is an endless Fibonacci loop โ it runs until the sequence overflows 16 bits and wraps, then keeps going.
LDI R1, 0 # R1 = 0 (current)
LDI R2, 1 # R2 = 1 (next)
NOP # loop start
ADD R3, R1, R2 # temp = R1 + R2
NOP
ADD R1, R2, R0 # R1 = R2 (R0 is zero, so this is a copy)
NOP
ADD R2, R3, R0 # R2 = temp
JMP 2 # back to loop startThe NOPs are there because the register file writes on the clock edge โ a result isn't readable by the very next instruction. Spacing the dependent instructions apart is the simplest way to respect that on a design with no forwarding or interlocks. Building a CPU that can issue those instructions back to back means adding a hazard unit, and that's the next revision.
Each module was simulated in Icarus Verilog and inspected waveform by waveform in GTKWave before being wired into the top level. The load and store instructions were the last to work โ the store encoding uses rd as the address base rather than the destination, which is easy to get backwards and produces silently wrong memory writes rather than an obvious failure.
After simulation passed, the design was synthesized for a Tang Nano 9K FPGA. The board's six LEDs are driven by the low six bits of the program counter, so the loop is visible as a counting pattern โ real hardware executing the instruction set, not a testbench asserting that it should.
View the Verilog on GitHubI'm Doruk Orak, a 12th grade student at Sankt Georg Austrian High School in Istanbul.
My interests span hardware design, artificial intelligence, programming, and physics. Connecting the world of intelligent systems with physical circuits.