Tao On Twitter





    Showing posts with label ASIC. Show all posts
    Showing posts with label ASIC. Show all posts

    Saturday, September 27, 2008

    Silicon Biometrics : The OCV Authentication Solution

    I never thought I'd say this ever: There's finally a reason to appreciate OCV and we can all thank Verayo for it. In one of my earlier posts, I wrote about the growing problem of counterfeit ASICs (Will The Real ASIC Please Stand Up?)The variation of on-die device and metal parameters that is the bane of designers all over the globe has turned out to be useful in the most unique way. OCV, it seems, is the silicon equivalent of a fingerprint. Let us review the salient points about OCV:

    • On-die variations are always present in all devices
    • Although on-die variations obey statistics, individual variations are essentially random
    • The OCV solution space is huge. Assume that a single transitor can be one of a types and a single net can be any one of b types. If you have x transistors and n nets, a single ASIC can be any one of a *b * c *d possibilities.
    • The probability of two ASICs having the exact same characteristics is so small, it is essentially zero.
    The existence of OCV is half the solution. It sort of like saying I can uniquely identify a grain of sand by the fact that each of its trillion odd molecules is differently positioned when compared to any other grain of sand. For the solution to be complete, there has to be a feasible way of measuring OCV for a given device or its effects. Hmm, Why do I worry about OCV? what's the worst that can happen? I'm guessing the phrase "timing violations" are flashing in six foot neon letters in your head right about now.

    By creating a circuit with lots of reconvergent logic with very low (or zero margins) margin setup and hold paths, you can observe the effects of OCV. The values captured at the output will not only change with each device but also change with the input stimulus. Depending on the stimulus, the path taken through the reconvergent cone may or may not suffer from a timing violation. By observing the outputs for a set of random stimuli, each device can be uniquely identified.

    If you want to know more about this technology, it's official name is "Physically Unclonable Functions(PUF)" . For the academically oriented, the home page of Professor Srini Devadas (MIT/ Verayo Co-founder) has download links for all his published PUF papers.

    Tags : ,

    Sunday, August 17, 2008

    The Hidden Factory : First Time Yield In The ASIC Design Flow

    The Six Sigma methodology is widely used to redesign processes to be more efficient. The way Six Sigma works is to redesign a process to all but eliminate defective outputs. To achieve Six Sigma, a process must create no more than 3.4 defects per million opportunities. Why the focus on defects?

    • Defects = Waste: When a process creates a defective output, all the effort and material invested in that defective part is essentially wasted.
    • Defects = Rework: When an intermediate stage produces a defective output, the process for that stage is repeated to produce a correct output.
    Focusing on eliminating defects in a systematic manner allows a process to be both more efficient as well as produce outputs of consistently high quality.

    One of the key measures of process quality used in Six Sigma is First Time Yield(FTY). The first time yield of a single stage is the probability that a correct output is created if the stage is run exactly once. The first time yield of a process is the product of the first time yield of its stages.

    FTY is great for identifying priority areas for redesign. Think of a simple process consisting of two stages: A and B. The FTY of A is 50%. The FTY of B is 100%. The FTY of the entire process is 50% (100% * 50% ). Stage B is perfect but Stage A is bringing down the FTY of the process. Another benefit of FTY is to identify "hidden factories". What if stage A is followed by a quality check stage that mandates a rerun of A in case of defective outputs? If we were to insert a QC check stage between A and B, the process as a whole will have 100% yield. But, stage A will be repeated twice on average to produce a correct output for a given input. When you view the process as a black box, you would not see these stage A iterations. For this reason, these iterative loops are called hidden factories. So, a process of multiple stages can produce a million correct outputs for a million correct inputs and still not be a good process.

    Consider your ASIC Design Flow in this context:
    • What are the chances that a design will go through your ASIC Design Flow in one shot?
    • What are the chances of a particular flow (synthesis, scan insertion,...) will go through in one shot?
    • Which flow is bringing you down?
    • Where are the hidden factories?
    • What are you going to do about it?
    Tags : ,

    Thursday, June 19, 2008

    Out Now! : TSMC Reference Flow 9.0 Is Now Available

    The TSMC Reference Flow 9.0 is available for download from TSMC-Online. Eyecatching items include:

    • DFT
      • Using E-fuse for MBIST
      • Failure Analysis
      • Low-power ATPG
    • Really Advanced CTS
      • CTS for Dynamic IR
      • CTS for Low-power
      • Multi-Mode Multi-corner CTS
    • Low-power
      • Low power automation with UPF
    • Statistical Design
      • {LPC, CAA, VCMP} --> {Timing, Power, Leakage} Flows

    Tags : ,

    Thursday, June 12, 2008

    SNUG 2008 : Registrations Open

    In case you're a Synopsys customer in Bangalore, registration for SNUG2008 is now open. Why, Aditya, thank you for that perfectly selfless propagation of useful information with no ulterior motives....

    NOT!

    If you can, do try and attend my presentation (in the Synthesis & Test track) on the 10th of July.

    Register Cloning For Accelerated Design Closure

    Multiple technologies exist to achieve timing closure on critical paths. One such technology, clock skew optimization, changes the arrival of clock edges at the launch and sink registers to increase the effective clock period of the critical path. Standard clock skew optimization does not necessarily utilize the full slack available at the input of a register but only the amount required to resolve the setup violations on paths from the register. If clock skew optimization were to utilize the input slack to the fullest extent towards have a large setup slack on the erstwhile critical path, it could accelerate setup timing closure by letting the tool concentrate on other paths in the design. However, the process could also introduce a large number of hold violations on other paths from the register with low hold slack due to the early launch of data. By having separate clone registers for setup and hold paths, one can fully utilize the input slack to launch registers for accelerating timing closure while limiting resultant hold violations. Since cloned registers are exact copies of the original register, the impact of register cloning on verification and ECO methodology effort is minimized. In this paper, a methodology will be presented to identify cloning candidates, insert clone registers and verify the final design against the un-cloned input.


    Tags : ,

    Sunday, May 04, 2008

    Force Multipliers : A Paradigm for Technology and Tool Development

    The concept of force multipliers originated in military science. Force multipliers are fulcrums that allow you to have an effect that is exponentially proportional to the amount of resources employed. Radar is a good example of a force multiplier. An air force that uses radar will be able to successfully attack or fend off a much larger force that does not have the benefit of the technology simply by being able to track their opponents in the battlefield. Sometimes, the force multiplier does not physically exist. For example, the coordinated use of air and ground forces as a blitzkrieg is more effective than an uncoordinated force of the same size. The blitzkrieg tactic is the force multiplier in this case.

    It seems to me that viewing tools and technologies as force multipliers is a good paradigm for guiding technology and tool development.

    1. Force multipliers do not replace; they complement. Did they disband the air force after inventing radar? The objective of tool development in a force multiplier context should not be to create a software version of your engineer. It's about creating a tool that will allow them to control a lot with very little. In a design environment such as Pyramid or IC-Catalyst, the engineer is in control but the environment takes care of all the small stuff (generating scripts, checking reports, firing jobs...). An engineer using a design environment is in a position to accomplish a lot with very little effort.
    2. Force multipliers need not be complex; just effective. Creating and supporting a complete design environment is a lot of work. plus, there's always the chance that you're over-solving the problem. Small utilities addressing the right issues in the flow can add value with much less effort. Think scalpels, not broadswords.
    3. Force multipliers need not physically exist.A stable design methodology does not physically exist, and yet, guides the engineer in the direction that will produce the optimal result in the shortest time.
    4. Force multipliers stack. The effect of more than one force multipliers is not the sum of their individual effects but their product. For example, a good hierarchical methodology and a good timing fix utility employed together will accelerate timing closure beyond the sum of their individual effects.

    Tags : ,

    Saturday, February 23, 2008

    Will The Real ASIC Please Stand Up? : The Brave New World Of Counterfeit ASICs

    Just last week, Reuters reported that counterfeit components worth $1.3 Billion were seized in a joint operation by the US and the EU. Chew on these stats:

    • These counterfeits must be pretty sophisticated. Biggies like Intel, Phillips and Cisco are not exactly known for making great op-amps.
    • Counterfeiters are going after the big-ticket items. If 360,000 parts were seized, that puts the average value of the components seized at $3600!
    Counterfeit ASICs bring us face to face with an altered reality. You can fake a watch, a perfume or even clothing, but an ASIC? Counterfeits used to be something Nike and Armani worry about, not ASIC design companies. Clearly, we're not in Kansas anymore. The big questions are:
    • Do these counterfeits actually work??
    • How are ASICs reverse-engineered?
      • Do they use the datasheet/spec?
      • Do they obtain the GDSII?
      • Do they strip the die layer by layer?
    • Can we prevent an ASIC from being faked?
      • If the datasheet and the chip are out in the real world, can we prevent a copy?
    • Can we authenticate an ASIC beyond doubt?
      • What prevents them from copying that, too?
    Tags : , ,

    Thursday, February 07, 2008

    I Coulda Been A Contender : Some VLSI Ideas I Wish I Had Had (First)

    Getting out plan for the new year got me thinking about issues in the ASIC flow and what could (or should) be developed to make the flow better. It could be a new technology, a new tool or a new flow even. The hardest part of this is to be able to break out of the box and approach issues from a new angle. Naturally, at such times, I look back at some ideas I heard about and think "I wish I'd thought of that!" It's not that these ideas made their inventors rich but they did make me sit up and take notice because of their refreshingly different viewpoint. Here's hoping they make you think, too.

    • Wire Tapering : Minimize the delay of a wire by controlling its shape. Rather than the usual rectangular shape, the optimal shape (in terms of delay) for an interconnect is exponentially tapered. The wire will be thickest near the driver and thinnest near the sink. In a regular net, the capacitance of the end section of the wire is the same as the initial section of the wire. Since the end capacitance has to be driven through a large resistance (basically the resistance of the whole net), the driver "sees" this end capacitance as a large load. By tapering the interconnect, the capacitative load decreases with distance from the driver. Thus, the overall load seen by the driver for a tapered net is lesser than a rectangular one.
    • Configurable Processors : Adapt a processor's instructions to match the application. Most processors are jacks of all trades but masters of none. Though processor cores are not particularly good at anything, using them for implementing features has its advantages. There's low risk of functional bugs within the core itself. Fixing bugs and adding features is as easy as updating the firmware. This is the low-performance/low-effort solution. Custom RTL, on the other hand, is the high-performance/high-effort solution. Implemented features have high performance but verification requires a lot of effort and bug fixes and additional features will atleast require a respin. Configurable processors allow you to get the best of both worlds. By implementing an instruction set tailored to the end application, the performance of the core is increased while verification and feature updates and bug fixes are still relatively easy.
    • Channel-less Floorplan: Why not just route top-level signals through the block? Look at the channels in your hierarchical floorplan (the gaps between the blocks). You need these channels to be able to route top-level signals from block to block. Usually, these signals don't undergo logical transformations at the top-level so it's pretty much getting the signal from point A to point B using repeaters. In a channel-less floorplan, the repeaters don't go round or over the block; They go through the block. The idea is to create feedthroughs in blocks such that a signal can use these feedthroughs to get to the other side of the block faster. The benefit is that you save on the die area that was previously dedicated to channels.
    • Asynchronous ASIC Design: What if there were no clocks? Clocks are nothing but synchronization signals. At each clock edge, you are guaranteed that the input data is stable and valid. The problem with having a clock is that your design can run only as fast as the worst path in the design. Even if 99.99999% of paths run at 1ns , the last path running at 2ns requires you to run the entire design at 500MHz. Wouldn't it be great if the performance of the entire device did not depend on the worst path in the design? That's where asynchronous designs come in. Asynchronous ASIC designs do not require a clock. Instead of clock-based synchronization, these designs use handshaking, semaphores and other methods to exchange data. Each part of the device runs as fast as it can and the effective frequency of operation is no longer determined by the worst path in the design. Other benefits of asynchronous design include lesser power dissipation (no clock trees) and lesser dynamic IR drop issues (without clocks, transitions in the design are randomized).
    • Analog Computing: Forget binary and use (semiconductor) physics. There's something unnatural about digital design. In a world where everything sentient is analog, should computing be any different? The idea behind analog computing is to use physical laws for computation. Suppose you want to build an adder, currents equivalent to the numbers to be added are passed through a resistance. The voltage drop across the resistance is proportional to the sum. The square-law characteristic of the MOS transistor in saturation is used to create computation devices such as multipliers. By piggy-backing onto the mathematics of physical effects, it is possible to use analog for computation at a fraction of the delay, area and power of digital circuits.

    Tags : ,

    Wednesday, November 14, 2007

    The ASIC Factory (Part 1) : The Toyota Way For Fabless ASICs

    Reading The Toyota Way by Jeffrey Liker got me to thinking about the benefits of bringing manufacturing into the realm of ASIC design (low cost, high quality, predictability, etc). For those who don't know, The Toyota Way is Toyota's management philosophy. The production system reflecting that philosophy allows Toyota to consistently figure as one of the best companies in the world. It's hard to comprehensively describe a philosophy in words but it is generally accepted that there are 14 principles that capture the essence of the Toyota Way. How can these 14 be applied to ASIC design? My thoughts:

    #1. Base your management decisions on a long-term philosophy, even at the expense of short-term financial goals.

    Focus on core competencies and don't waste too much energy pursuing multiple courses of action. This would apply to designs too. Trying to be good at everything from low-power wireless designs to high-performance multi-core processors is a recipe for disaster. You could extend this to methodologies or even EDA tools. Streamline. Focus. On the flip side, don't let the lack of a large current market prevent you from pursuing technologies or products that would have a great future market. Lastly, distinguish between the two (easier said than done but someone has got to say it ;) ).

    #2. Create a continuous process flow to bring problems to the surface.

    Have a methodology and design process that is transparent and efficient. It will allow you to easily spot problems in the flow. Minimize idle time and non-value added work. In the course of work, a design engineer:

    • writes a script
    • checks the syntax
    • executes the script
    • waits for the results
    • opens some reports
    • checks specific parameters (slack, perhaps)
    Writing the script and checking the slack are pretty much the only steps that really adds value. Everything else is a waste of the engineer's time. What are you doing about it?

    Tags : , ,

    Tuesday, October 23, 2007

    Triple Digits in Four Years : Open-Silicon Books 100th Design Win

    I'm happy to say that Open-Silicon has now won a total of 100 designs in a short span of just 4 years. These design wins are across the spectrum in terms of both applications (low-power wireless to high-performance cluster nodes) and processes (0.25u to 45nm). Here's what Pierre Lamond (of Sequoia Capital) had to say about the company and its achievement:

    "Open-Silicon has been an agent for change in the ASIC market from the moment they launched their innovative business model. Their ability to book 100 design wins in four years is a testament to the market's need for highly predictable and reliable custom silicon. Open-Silicon's skill at delivering on their customer's critical time-to-market requirements is what makes them one of the fastest growing companies in the market."

    Tags : ,

    Monday, October 22, 2007

    It's Not What You've Got, It's How You Use It : Ideas For Cost-effective Server Farms


    The optimal utilization of resources is instrumental in improving the profitability of any enterprise. As a fabless ASIC company scales in size, the issue of automated/intelligent resource management becomes a bottleneck. You're too big to be able make snap decisions in response to immediate demand. There has to be a formal process of measuring resource requirements and subsequent acquisition of said resources. One such resource that requires intelligent estimation and acquisition are compute resources available to your engineers. The compute resources referred to are the sum total of compute power in the enterprise that are capable of executing EDA tools . These include both the dedicated high-end compute servers and often overlooked workstations that sit in every engineer's cubicle. It is true that machines will probably come third when it comes to expenses (after people and EDA licenses) but there is a lot of scope for more efficient machine usage because it is third in line! Some ideas to make the most of your servers:

    1. Get a queue. When resources are accessed through a load-sharing mechanism such as LSF or FlowTracer, it is a whole lot easier to analyze trends and increase utilization across all machines. Low-power desktops that do not usually perform anything more intensive than a screen refresh will, on the queue, be put to use on jobs that meet their memory limits and save your high-end servers for jobs that require them. A enterprise with a queue has the advantage of a scalable system. Your engineers see a single interface whether you have ten machines or a thousand.
    2. Choose Your Machines Carefully. With data on jobs and memory, one can can create a machine pool that reflects these requirements. Why is this important? Have a look at machine prices on Epinions. An 8GB single-CPU Opteron will cost you $900 a piece . On the other hand, a 64GB machine with 4 CPUS will cost you $17400. If most jobs take up less than 8GB of memory, you'd be better off choosing low-end servers instead of high-end ones. Why get a 64GB machine when you can get 19 8GB opterons instead?
    3. Spread it Out. Ensure that there's a mechanism to even out demand throughout the working day or week. One sign that you might need such a policy is that there don't seem to be enough machines during the day but most of your servers are idle at night! In the absence of such measures, you might end up with more machines that are acquired to meet just the peak load. By instituting a policy that ensures an even utilization across the day, you require lesser resources while utilizing those resources to the maximum.
    Tags : , , , , ,