ASIC Verification: Verilog
Showing posts with label Verilog. Show all posts
Showing posts with label Verilog. Show all posts

Thursday, September 4, 2008

Verilog questions

1. Is this a valid, synthesizable, use of a for loop?

module for_loop();

reg [8:0] A, B;
integer i;
parameter N=8;
always@(B)
begin
for (i=1; i<=N; i=i+1)
A[i-1]=B[i];
A[N] = A[N-1];
end
endmodule

2. Assuming the code above is synthesizable, Which of the following continuous assignment statements would have the closest meaning?

A. assign A = B << 1;
B. assign A = B <<< 1;
C. assign A = B >> 1;
D. assign A = B >>> 1;

3. If the following logic is built exactly as described, which test vector sensitizes a stuck-at-0 fault at "e" and propagates it to the output "g".

module (a, b, c, d, e, f, g);
input a, b, c, d;
output e, f, g;

assign e = a & b;
assign f = c ^ e;
assign g = d | f;

endmodule

A. {a, b, c, d} = 4’b0010;
B. {a, b, c, d} = 4’b1100;
C. {a, b, c, d} = 4’b1111;
D. {a, b, c, d} = 4’b0101;
E. None of the above

4. Consider the following two test fixtures.

// Fixture A
parameter delay1 =
parameter delay2 =
initial
begin
B = 1’b0;
#20 A = 1’b1;
#delay1 A = 1’b0;
#delay2 B = 1'b1;
end

// Fixture B
initial
fork
B = 1’b0;
#20 A = 1’b1;
#40 A = 1’b0;
#60 B = 1’b1;
join

For these two fixtures to produce the same waveforms, delay1 and delay2 have to be set
as follows:

A. delay1 = 40; delay2 = 60;
B. delay1 = 30; delay2 = 20;
C. delay1 = 30; delay2 = 30;
D. delay1 = 20; delay2 = 20;
E. None of these are correct

5. In verification, most of the effort should be applied at the system (complete chip) level. Which of the following statements gives the best reason as to why?

A. Most of the bugs in a design are in the netlist wiring it together.
B. Most of the bugs in a design are due to poorly understood interactions between different modules.
C. This is the fastest way to verify the individual modules that make up the design.
D. Most of the bugs in a design occur because of poorly designed interfaces, e.g. buses.
E. None of the above are remotely a good reason.

6. Consider the following specify block:

specify
specparam A0spec = 1 : 2 : 3;
specparam A1spec = 2 : 3 : 4;
(a => b) = (A0spec, A1spec);
endspecify

This is defining the following:

A. Rising, falling and steady delay from input a to output b of 1, 2, and 3 ns respectively when a is 0, and 2, 3, and 4 ns when a is 1.
B. Minimum, typical and maximum delay from input a to output b of 1, 2, and 3 ns on a rising edge at B, and 2, 3 and 4 ns on a falling edge.
C. Non-blocking assignment of a to b with minimum, typical and maximum delay of 1, 2 and 3ns.
D. Setup time requirements for the flip-flop with output B.



Wednesday, July 23, 2008

Gray Code Counter Implementation

A Gray code is an encoding of numbers so that adjacent numbers have a single digit differing by 1. The term Gray code is often used to refer to a Binary Reflected Gray Code. We can implement a gray code counter in a different ways. Consider the following table carefully.

B : 000, 001, 010, 011, 100, 101, 110, 111
G: 000, 001, 011, 010, 110, 111, 101, 100

To convert a binary number d1,d2,..,d(n-1),dn to its corresponding Binary Reflected Gray Code, start at the right with the digit dn (the LSB). If the d(n-1) is 1, replace dn by (1-dn); otherwise, leave it unchanged. Then proceed to d(n-1). Continue up to the first d1, which is kept the same. The resulting number g1,g2,..,g(n-1),gn is the Reflected Binary Gray Code.

The most common Gray code is where the lower half of the sequence is exactly the mirror image of first half with only the MSB inverted. We illustrate the 3-bit binary Gray code as an example.

Binary to gray code can be achieved by

gray[2] = binary[2];

gray[1] = binary[2] ^ binary[1];

gray[0] = binary[1] ^ binary[0];

A simple verilog code to implement this function is given by

assign gray = (binary>> 1) ^ binary; // Right shift by 1 and EX-OR with binary.

 module gray_cntr (  
clock_in,
rst_n,
enable_in,
cnt_out
);


// I/O Declarations

input clock_in, rst_n, enable_in;
output [ 2:0] cnt_out;
wire [2:0] cnt_out;
reg [2:0] cnt;

always @ (posedge clock_in or negedge rst_n)
if (!rst_n)
cnt
<= 1'b0;
else if (enable_in)
cnt
<= cnt + 1'b1;

assign cnt_out = { cnt[2], (^cnt[2:1]), (^cnt[1:0]) };

endmodule

Sunday, June 15, 2008

RTL Design techniques - Pre-RTL Checklist.

Your success in IC design is directly depends on your RTL code. There is a lot more that goes into a good RTL description than just writing with good coding style. Design for Test and Design for Synthesis are just a few examples of design goals that can be affected at the RTL. This post is all about RTL design issues. Code it correctly from the beginning and you won'’t need so many big fancy tools to solve your timing closure problems at the back end of the design cycle.

There are many design issues - which impact the speed and area of the design - need to be resolved before you begin coding your design.

Communicate design issues with your team - Things to be worked out as a team
  • Naming convention for hierarchical blocks,
  • Naming convention for signals,
  • Active low or active high states for the signal
Does the specification define how the design should be partitioned?
Partitioning helps to break down your big design into smaller blocks and assign each small unit to different members of the team. Follow the specification's recommendation for partitioning.

What are the I/O requirements?
At the major functional block level, define the interface protocol as soon as possible. What bus interface protocol will be used? PCI, AHB or OCP. Get the specification for each bus and interface to the design before you begin coding. Make sure the function and timing of each one is clear. This will also enable you to create high level models of your design before you start coding the RTL.

What about the clocks in the design?
How many clocks will be required for the design? Where are the clocks for the chip coming from? Will they be internally generated? PLL? Divide by circuits? Externally supplied clocks? You have to isolate your clock generation circuitry from the rest of the chip design. Especially if it is analog based.

What other IPs are you using?
Does the design require any extra IP (Intellectual Property) to be integrated into it? RAMs? Cores? Buses? FIFOs? Then start with the interface to each IP block and define it.

Is it your expectation that you are pin-limited or gate limited?
Being pin-limited means that you don’t have enough I/O pads in your ASIC package to do what you really want to do. You might be able to double up on the functions of each pin, which would require multiplexing signals and would prevent any ideas of a unidirectional bus interface at the I/O pad level. But if you need all the signals to be active simultaneously, you won’t be able to do it either. You'’ll have to split the design up. You should know before you begin your RTL.

Being gate-limited means that the design has too much functionality for the die size chosen. You might have to cut out functionality to fit on the die. Or you can try to optimize your design for area, which means speed objectives might be tough to meet. It is hard to estimate whether you will be gate limited at the beginning of a project unless you have been through this design before.

Is it your expectation that you will be pushing the speed envelope of the technology?
  • How much functionality are you putting into your design
  • At what speed will it be running?
  • What technology are you going to use to implement it?
  • Has it ever been done before?
  • What changes to the design are you willing to make to achieve the speed goal for your design? Pipelining or Register re-timing.
In the next post, I'm going to post the rules that tend to cause the most common errors.

Saturday, May 31, 2008

VHDL vs Verilog

What is the reason that Verilog is usually considered better at low level modeling than VHDL? Why is VHDL usually considered better than Verilog for high level modeling?

Verilog has built-in types for gates and transistors, can also handle true bidirectional signals (VHDL has none of these things).

VHDL allows users to define their own data types which allows users to extend the language. Also, support for libraries and packages lends itself to more complex models.

Tuesday, May 20, 2008

Asynchronous and Synchronous Reset

ASYNCHRONOUS RESET

A fully asynchronous reset is one that both asserts and de-asserts a flip-flop asynchronously. Here, asynchronous reset refers to the situation where the reset net is tied to the asynchronous reset pin of the flip-flop. Additionally, the reset assertion and de-assertion is performed without any knowledge of the clock. This type of reset is very common but is very dangerous if the module boundary represents the FPGA boundary.

The biggest problem with the asynchronous reset circuit described above is that, it will work most of the time. However, if the edge of the reset deassertion is too close to the clock edge and violate the reset recovery time, then the output of FF goes to metastable. The reset recovery time is a type of setup timing condition on a flip-flop that defines the minimum amount of time between the de-assertion of reset and the next rising clock edge as shown in Figure.


It is important to note that reset recovery time violations only occur on the de-assertion of reset and not the assertion. Therefore, fully asynchronous resets are not recommended.

SYNCHRONOUS RESET

The most obvious solution to the problem introduced in the preceding section is to fully synchronize the reset signal as you would any asynchronous signal.

The advantage to this type of topology is that the reset presented to all functional flip-flops is fully synchronous to the clock and will always meet the reset recovery time. The interesting thing about this reset topology is actually not the deassertion of reset for recovery time but rather the assertion In the previous section, it was noted that the assertion of reset is not of interest, but that is true only for asynchronous resets and not necessarily with synchronous resets. Consider the scenario illustrated in Figure.
Consider the scenario where the clock is running sufficiently slow, the reset is not captured due to the absence of a rising clock edge during the assertion of the reset signal. The result is that the flip-flops within this domain are never reset.

Fully synchronous resets may fail to capture the reset signal itself (failure of assertion) depending on the nature of the clock.

For this reason, fully synchronous resets are not recommended unless the capture of the reset signal (reset assertion) can be guaranteed by design.

Asynchronous Assertion, Synchronous De-assertion

A third approach that captures the best of both techniques is a method that asserts all resets asynchronously but de-asserts them synchronously.
In Figure, the registers in the reset circuit are asynchronously reset via the external signal, and all functional registers are reset at the same time. This occurs asynchronous with the clock, which does not need to be running at the time of the reset. When the external reset de-asserts, the clock local to that domain must toggle twice before the functional registers are taken out of reset. Note that the functional registers are taken out of reset only when the clock begins to toggle and is done so synchronously.

A reset circuit that asserts asynchronously and de-asserts synchronously generally provides a more reliable reset than fully synchronous or fully asynchronous resets.


The code for this synchronizer is shown below.

module reset_sync(
output reg rst_sync,
input clk, rst_n);
reg R1;
always @(posedge clk or negedge rst_n)
if(!rst_n)
begin

R1 <= 1'b0;
rst_sync <= 1'b0;
end

else
begin

R1 <= 1'b1;
rst_sync <= R1;
end
endmodule

Tuesday, May 13, 2008

VLSI FAQ

Hi All,
I am back with bank after my UCSD mid-term test. Today I am going to post some of the important verilog questions.

How to solve setup & Hold violations in the design?

To solve setup violation
  • Optimizing / Restructuring Combinational logic between FFs.
  • Tweak flops to offer lesser setup delay.
  • Tweak launch-flop to have better slew at the clock pin, this will make CK->Q of launch flop to be fast there by helping fixing setup violations.
  • Play with skew (Tweak clock network delay, slow-down clock to capturing flop and fasten the clock to launch-flop) (otherwise called as Useful-skews)
To solve Hold Violations
  • Adding delay / buffer [as buffer offers lesser delay, we go for special Delay cells whose functionality Y=A, but with more delay]
  • Also, one can add lockup-latches [in cases where the hold time requirement is very huge, basically to avoid data slip]
What is tie-high and tie-low cells and where it is used?

Tie-high and Tie-Low cells are used to connect the Gate of the transistor to either Power or Ground. In deep sub micron process, if the Gate is connected to Power / Ground, the transistor might be turned ON / OFF due to power or ground bounce. The suggestion from foundry is to use Tie cells for this purpose. These cells are part of standard-cell library. The cells which require Vdd, comes and connect to Tie high.(so tie high is a power supply cell), while the cells which wants Vss connects itself to Tie-low.

What is metastability and steps to prevent it?

Metastability is an unknown state - it is neither Zero nor One. Metastability happens for the design systems violating setup or hold time requirements. Setup time is a requirement, that the data has to be stable before the clock-edge and hold time is a requirement, that the data has to be stable after the clock-edge. The potential violation of the setup and hold violation can happen when the data is purely asynchronous and clocked synchronously.

Steps to prevent Metastability:
  • Using proper synchronizers(two-stage or three stage), as soon as the data is coming from the asynchronous domain.
  • Using Faster flip-flops (which has narrower Metastable Window).
What is local-skew, global-skew,useful-skew mean?

Local skew : The difference between the clock reaching at the launching flop vs the clock reaching at the destination flip-flop of a timing-path.
Global skew : The difference between the earliest reaching flip-flop and latest reaching flip-flop for a same clock-domain.
Useful skew: Useful skew is a concept of delaying the capturing flip-flop clock path, this approach helps in meeting setup requirement within the launch and capture timing path. But the hold-requirement has to be met for the design.

What are the various timing-paths which i should take care in my STA runs?
  1. Timing path starting from an Input-port and ending at the Output port (purely combinational path).
  2. Timing path starting from an Input-port and ending at the Register.
  3. Timing path starting from an Register and ending at the Output-port.
  4. Timing path starting from an Register and ending at the Register.

What are the various Design constraints used while performing Synthesis for a design?
  1. Create the clocks ( Frequency, Duty-Cycle).
  2. Define transition-time requirements for the input-ports
  3. Specify load values for the output ports
  4. For the inputs and the outputs, specify the delay values (input delay and ouput delay), which are already consumed by the neighbour chip.
  5. Specify the case-setting (in case of a mux) to report the timing to a specific paths.
  6. Specify the False-paths in the design
  7. Specify the Multi-cycle paths in the design.
  8. Specify the clock-uncertainity values (with respect to jitter and the margin values for setup/hold).
  9. Specify few verilog constructs which are not supported by the synthesis tool.
What is meant by wire-load model?

In the synthesis tool, in order to model the wires we use a concept called Wireload models. Wireload models are statistical based on models with respect to Fanout. Say, for a particular technology based on our previous chip experience we have a rough estimate we know if a wire goes for "n" number of fanout, then we estimate its delay as say "x" delay units. So a model file is created with the fanout numbers and corresponding estimated delay values. This file is used while performing Synthesis to estimate the delay for Wires and to estimate the delay for cells.

What are the measures or precautions to be taken in the Design when the chip has both analog and digital portions?

As today's IC has analog components also inbuilt, some design practices are required for optimal integration. Ensure in the floor-planning stage that the analog block and the digital block are not siting close-by, to reduce the noise. Ensure that there exists separate ground for digital and analog ground to reduce the noise. Place appropriate guard-rings around the analog-macro's. Incorporating in-built DAC-ADC converters, allows us to test the analog portion using digital testers in an analog loop-back fashion. Perform techniques like clock-dithering for the digital portion.

What is meant by inferring latches, how to avoid it?

Consider the following :

always @(s1 or s0 or i0 or i1 or i2 or i3)
case ({s1, s0})
2'd0 : out = i0;

2'd1 : out = i1;
2'd2 : out = i2;
endcase

In a case statement if all the possible combinations are not compared and default is also not specified like in example above, a latch will be inferred. In above case if {s1,s0} =3, the previous stored value is reproduced. The same may be observed in IF statement in case an ELSE IF is not specified. To avoid inferring latches make sure that all the cases are mentioned if not default condition is provided.

Tell me structure of Verilog code you follow?

A good template for your Verilog file is shown below.

// Timescale directive tells the Simulator the Base units and Precision time unit of the simulation
`timescale 1 ns / 10 ps
module name (input and outputs);
// Parameter Declarations
parameter parameter_name = parameter value;
// Input / Output Declarations
input in1;
input in2; // Single bit Inputs
output [msb:lsb] out; // A Bus Output
// Internal signal register type declaration -
// Register types (only assigned within always statements).
reg register variable 1;

reg [msb:lsb] register variable 2;
// Internal signal. net type declaration - (only assigned outside always statements)
wire net variable 1;

// Hierarchy - Instantiating another module

reference name instance name (
.pin1 (net1),
.pin2 (net2),
.
.pinn (netn)
);

// Synchronous Procedures

always @ (posedge clock)
begin
.
end

// Combinatinal Procedures

always @ (signal1 or signal2 or signal3)
begin
.
end

assign net variable = combinational logic;


endmodule


Monday, March 31, 2008

Verilog FAQ1

What are the ways to create a race condition and how can these race conditions can be avoided?

The IEEE Verilog Standard defines which statements have a guaranteed order of execution and which statements don' t have a guaranteed order of execution.

A Verilog race condition occurs when two or more statements that are scheduled to execute in the same simulation time-step, would give different results when the order of statement execution is changed.

module race (out1, out2, clk, rst);
output out1, out2;
input clk, rst;
reg out1, out2;
always @(posedge clk or posedge rst)
if (rst) out1 = 0;
else out1 = out2;

always @(posedge clk or posedge rst)
if (rst) out2 = 1;
else out2 = out1;
endmodule

If the first always block executes first after a reset, both out1 and out2 will take on the value of 1. If the second always block executes first after a reset, both out1 and out2 will take on the value 0. This clearly represents a Verilog race condition.

Making multiple assignments to the same variable from more than one always block is a Verilog race condition, even when using nonblocking assignments.

One of the recommendations is to avoid driving variables from multiple sources.

Illustrate example of how unintentional deadlocked situations can happen during simulation.

The deadlock situation is one in which one process is waiting for the other process to enable it, which in turn will enable the source process. The code could be a syntactically correct implementation, and still have a deadlock situation. The scenario can happen in both synchronous and asynchronous designs. A simple asynchronous example has been illustrated in the following, to demonstrate how deadlock occurs.

module deadlock;
reg reg1, reg2;

initial
begin
reg1 = 1'b0;
wait @ (reg2==1'b1)
end

always @(reg1)
begin
if (reg1==1'b1)
reg2 = 1'b1;
end
endmodule

The above example is an illustration of the deadlock scenario, which can be difficult to capture in a larger implementation.

What is the difference between a vectored and a scalared net?

Both scalared and vectored are Verilog constructs used on multi-bit nets to specify whether or not specifying bit and part select of the nets is permitted. For example,

wire scalared [3:0] a;
wire vectored [3:0] b;

wire c, d;

// Syntax error to use a bit select of vectored net
assign b[1] = 1'b1;
// OK
assign a[1] = 1'b0;

Difference b/w assign, de-assign and force, release.

The assign-deassign and force-release constructs in Verilog have similar effects, but differ in the fact that force-release can be applicable to nets, whereas assign-deassign is applicable only to registers.

The procedural assign-deassign construct is intended to be used for modeling hardware behavior, but the construct is not synthesizable by most logic synthesis tools. The force-release construct is intended for design verification, and is not synthesizable.

What does it mean to “short-circuit” the evaluation of an expression?

Verilog supports numerous operators that have rules of associativity and precedence. In some of the expressions, the result of the expression can be evaluated early on, due to the precedence and influence to override the rest of the expression. In that case, the entire expression need not be evaluated. This is called short-circuiting and expression evaluation.

For example,

assign out = ((a>b) & (c|d));

If the result of (a>b) is false (1'b0), then tools can already determine that the result of the AND operation will be 0. Thus, there is no need to evaluate (c|d) and rest of the equation is short-circuited.


What are the pros and cons of using hierarchical names to refer to Verilog objects?

The top-level module is called the root module because it is not instantiated anywhere. It is the starting point. To assign a unique name to an identifier, start from the top-level module and trace the path along the design hierarchy to the desired identifier.

assign status = top.hub_top.hpie.status_reg;

Adv:
  • It is easy to debug the internal signals of a design, especially if they are not a part of the top level pin out.
Disadv:
  • Sometimes, during synthesis, these hierarchical names get renamed, depending upon the synthesis strategy and switches used, and hence, will cease to exist. In that case, special switches need to be added to the synthesis compiler commands, in order to maintain the hierarchical naming.
  • If the Verilog code needs to be translated into VHDL, the hierarchical names are not translatable.
Does Verilog support an "a to the power b" operator?

Yes. Verilog supports the operation by using two astrices, back to back like,

assign out = (in ** 5);


Saturday, March 29, 2008

Verilog FAQ

What are the differences between blocking and nonblocking assignments?

There is one good paper by Stuart Sutherland about the blocking and non-blocking assignments. This paper can be downloaded from here.


Can you use a Verilog function to define the width of a multi-bit port, wire, or reg type?

The width elements of port declarations require a constant in both MSB and LSB. Before Verilog 2001, it is a syntax error to specify a function call to evaluate the value of these widths. For example, the following code is erroneous before Verilog 2001 version.

reg [ high(val1, val2) : low(val3, value4)] reg1;

In the above example, high and low are both function calls of evaluating a constant result for MSB and LSB respectively. However, Verilog-2001 allows the use of a function call to evaluate the MSB or LSB of a width declaration.

What is the difference b/w the following 2 lines of code?
#5 reg_a = reg_b;
reg_a = #5 reg_b;
Ans: 

Which one is better, asynchronous or synchronous reset for the storage elements?

There is one good paper by Stuart Sutherland about the synchronous and asynchronous reset. This paper can be downloaded from here.

What is the difference b/w the following 2 verilog codes?
a. assign c = condition ? a :b;
b. if(condition) c = a; else c = b;

Ans:


What logic gets synthesized when I use an integer instead of a reg variable as a storage element? Is use of integer recommended?

An integer can take the place of a reg as a storage element. The default width of the integer declaration is 32 bits. If you use integer in your RTL and store a 4 bit value, then the most significant 28 bits will be removed by the optimizer in the synthesis tool in order to minimize the area.

Although the use of integer is a legal construct, it is not recommended for the synthesis of storage elements.


How do you choose between a case statement and a multi-way if-else statement?

A case statement is typically chosen for the following scenarios:
  • When the conditionals are mutually exclusive and only one variable controls the flow in the case statement. The case variable itself could be a concatenation of different signals.
  • To specify the various state transitions of a finite state machine
  • Use of casex and casez allows use of x and z to represent don’t-care bits in the control expression
A multi way if statement is typically chosen in the following scenarios:
  • Synthesizing priority encoded logic
  • When the conditionals are not mutually exclusive
What is the difference between full_case and parallel_case synthesis directive?

There is one good paper by Stuart Sutherland about the full and partial case. This paper can be downloaded from here.

What is the difference b/w casex and casez statements? Which one is preferred?

Ans:
What is delta simulation time?
Ans:

How can you reliably convey control information across clock domains?

The readers are encouraged to read about good design implementation article here.


What are combinatorial timing loops? Why should they be avoided?

Combinatorial timing loops are hardware loops in which the output of either a gate or a long combinatorial path is fed back as an input to the same gate or to another gate earlier in the combinatorial path. These paths are generally created unintentionally when a variable from one combinatorial block is used to drive a signal that is used in the same combinatorial block from which the variable was derived. This typically happens in large size combinatorial blocks, wherein it is difficult to visually track that a loop is getting created.

These combinatorial feedback loops are undesirable for the following reasons:
  • Since there is no clock edge in between to break the path, the combinatorial loops will infinitely keep oscillating and triggering a square waveform, whose duty cycle is dependent upon the sum of ON delays and OFF delays across the combinatorial path.
Combinatorial loops can be caught quite early by one of the following means:
  • Periodic use of linting tools throughout the development process. This is by far the best and easiest way to catch and fix loops early in the design cycle.
  • If the loop is undetected during simulation, many synthesis tools have suitable reporting commands, which detect the presence of a loop. Note that synthesis tools proceed with the static timing analysis by breaking the timing arc of the loop for critical path analysis.