NVlabs
diff --git a/‎Dockerfile
Lines changed: 19 additions & 0 deletions b/‎Dockerfile
Lines changed: 19 additions & 0 deletions
diff --git a/‎LICENSE
Lines changed: 27 additions & 0 deletions b/‎LICENSE
Lines changed: 27 additions & 0 deletions
diff --git a/‎README.md
Lines changed: 65 additions & 54 deletions b/‎README.md
Lines changed: 65 additions & 54 deletions
diff --git a/‎data/VerilogEval_Human.jsonl
Lines changed: 156 additions & 0 deletions b/‎data/VerilogEval_Human.jsonl
Lines changed: 156 additions & 0 deletions
@@ -0,0 +1,19 @@
+FROM nvcr.io/nvidia/pytorch:22.08-py3
+LABEL maintainer="Mingjie Liu <mingjiel@nvidia.com>"
+RUN echo "alias python=python3" >> ~/.bashrc \
+        && echo "alias pip=pip3" >> ~/.bashrc
+RUN apt-get -y update \
+        && apt-get -y install vim 
+RUN apt-get install wget
+RUN apt-get install -y autoconf gperf flex bison screen
+RUN python -m pip install --upgrade pip
+RUN python -m pip install deepspeed scikit-learn pandas numpy scipy wandb
+RUN python -m pip install accelerate>=0.12.0 torch>=1.3 datasets>=1.8.0 sentencepiece!=0.1.92 protobuf evaluate
+RUN python -m pip install git+https://github.com/huggingface/transformers/
+RUN git clone https://github.com/steveicarus/iverilog.git && cd iverilog \
+        && git checkout 01441687235135d1c12eeef920f75d97995da333 \
+        && sh ./autoconf.sh && ./configure && make -j4\
+        && make install
+RUN python -m pip install jupyterlab
+RUN python -m pip install openai tiktoken
+ENV SHELL=/bin/bash
@@ -1,3 +1,28 @@
+MIT License
+
+Copyright (c) 2023 NVIDIA Research Projects
+
+Permission is hereby granted, free of charge, to any person obtaining a copy
+of this software and associated documentation files (the "Software"), to deal
+in the Software without restriction, including without limitation the rights
+to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
+copies of the Software, and to permit persons to whom the Software is
+furnished to do so, subject to the following conditions:
+
+The above copyright notice and this permission notice shall be included in all
+copies or substantial portions of the Software.
+
+THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
+IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
+FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
+AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
+LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
+OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
+SOFTWARE.
+
+
+This project contains code from human-eval (https://github.com/openai/human-eval/).
+
 The MIT License
 
 Copyright (c) OpenAI (https://openai.com)
@@ -19,3 +44,5 @@ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
 LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
 OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN
 THE SOFTWARE.
+
+
@@ -1,28 +1,48 @@
-# HumanEval: Hand-Written Evaluation Set 
+# VerilogEval: Evaluating Large Language Models for Verilog Code Generation 
 
-This is an evaluation harness for the HumanEval problem solving dataset
-described in the paper "[Evaluating Large Language Models Trained on
-Code](https://arxiv.org/abs/2107.03374)".
+This is an evaluation harness for the VerilogEval problem solving dataset
+described in the paper "[VerilogEval: Evaluating Large
+Language Models for Verilog Code Generation](https://arxiv.org/abs/2309.07544)".
+
+This evaluation dataset consists of 156 problems from the Verilog 
+instructional website [HDLBits](https://hdlbits.01xz.net/wiki/Problem_sets).
+We provide two sets of problem descriptions: machine generated and manually
+converted to text-only format.
 
 ## Installation
 
+We closely follow guidance from [HumanEval](https://github.com/openai/human-eval/tree/master).
+
 Make sure to use python 3.7 or later:
 ```
 $ conda create -n codex python=3.7
 $ conda activate codex
 ```
 
+Install [ICARUS Verilog](https://github.com/steveicarus/iverilog):
+```
+$ git clone https://github.com/steveicarus/iverilog.git && cd iverilog \
+        && git checkout 01441687235135d1c12eeef920f75d97995da333 \
+        && sh ./autoconf.sh && ./configure && make -j4\
+        && make install
+```
+
+It is recommended to use the provided [Dockerfile](https://github.com/NVlabs/verilog-eval/Dockerfile) 
+which already pre-installed ICARUS Verilog Simulator. Using the docker container
+you would still need to complete the following step.
+
 Check out and install this repository:
 ```
-$ git clone https://github.com/openai/human-eval
-$ pip install -e human-eval
+$ git clone https://github.com/NVlabs/verilog-eval
+$ pip install -e verilog-eval
 ```
 
 ## Usage
 
-**This program exists to run untrusted model-generated code. Users are strongly
+**This program would make system calls to *iverilog* and *vvp* to simulate 
+untrusted model-generated code. Users are strongly
 encouraged not to do so outside of a robust security sandbox. The [execution
-call](https://github.com/openai/human-eval/blob/master/human_eval/execution.py#L48-L58)
+call](https://github.com/NVlabs/verilog-eval/blob/main/verilog_eval/execution.py#L79-L112)
 in `execution.py` is deliberately commented out to ensure users read this
 disclaimer before running code in a potentially unsafe manner. See the comment in
 `execution.py` for more information and instructions.**
@@ -31,54 +51,46 @@ After following the above instructions to enable execution, generate samples
 and save them in the following JSON Lines (jsonl) format, where each sample is
 formatted into a single line like so:
 ```
-{"task_id": "Corresponding HumanEval task ID", "completion": "Completion only without the prompt"}
-```
-We provide `example_problem.jsonl` and `example_solutions.jsonl` under `data`
-to illustrate the format and help with debugging.
-
-Here is nearly functional example code (you just have to provide
-`generate_one_completion` to make it work) that saves generated completions to
-`samples.jsonl`.
-```
-from human_eval.data import write_jsonl, read_problems
-
-problems = read_problems()
-
-num_samples_per_task = 200
-samples = [
-    dict(task_id=task_id, completion=generate_one_completion(problems[task_id]["prompt"]))
-    for task_id in problems
-    for _ in range(num_samples_per_task)
-]
-write_jsonl("samples.jsonl", samples)
+{"task_id": "Corresponding VerilogEval task ID", "completion": "Completion only without the prompt"}
 ```
+We provide examples under `data/example` to illustrate the format and help with debugging.
 
 To evaluate the samples, run
 ```
-$ evaluate_functional_correctness samples.jsonl
+$ evaluate_functional_correctness samples.jsonl --problem_file data/VerilogEval_Human.jsonl
 Reading samples...
-32800it [00:01, 23787.50it/s]
+3120it [00:00, 16077.44it/s]
 Running test suites...
-100%|...| 32800/32800 [16:11<00:00, 33.76it/s]
+100%|...| 3120/3120 [00:32<00:00, 97.47it/s]
+Killing all hanging simulation process.
 Writing results to samples.jsonl_results.jsonl...
-100%|...| 32800/32800 [00:00<00:00, 42876.84it/s]
-{'pass@1': ..., 'pass@10': ..., 'pass@100': ...}
+100%|...| 3120/3120 [00:00<00:00, 30608.13it/s]
+{'pass@1': ..., 'pass@5': ..., 'pass@10': ...}
 ```
+
+The user must specify `--problem_file` input argument. We provide two sets of problem
+evaluations `data/VerilogEval_Machine.jsonl` and `data/VerilogEval_Human.jsonl`. 
+We also provide problem description files used to sample Verilog code completions 
+in `descriptions` directory.
+
 This script provides more fine-grained information in a new file ending in
 `<input_path>_results.jsonl`. Each row now contains whether the completion
 `passed` along with the execution `result` which is one of "passed", "timed
 out", or "failed".
 
-As a quick sanity-check, the example samples should yield 0.5 pass@1.
+As a quick sanity-check, the example samples should yield 0.5 pass@1. The results can be
+verified against the provided output 
+in `data/example/ExampleSolution.jsonl_reference.jsonl`.
 ```
-$ evaluate_functional_correctness data/example_samples.jsonl --problem_file=data/example_problem.jsonl
+$ evaluate_functional_correctness data/example/ExampleSolution.jsonl --problem_file=data/example/ExampleEval.jsonl
 Reading samples...
-6it [00:00, 3397.11it/s]
+6it [00:00, 221.60it/s]
 Running example suites...
-100%|...| 6/6 [00:03<00:00,  1.96it/s]
-Writing results to data/example_samples.jsonl_results.jsonl...
-100%|...| 6/6 [00:00<00:00, 6148.50it/s]
-{'pass@1': 0.4999999999999999}
+100%|...| 6/6 [00:00<00:00, 142.09it/s]
+Killing all hanging simulation process.
+Writing results to data/example/ExampleSolution.jsonl_results.jsonl...
+100%|...| 6/6 [00:00<00:00, 19941.22it/s]
+{'pass@1': 0.5}
 ```
 
 Because there is no unbiased way of estimating pass@k when there are fewer
@@ -90,26 +102,25 @@ $ evaluate_functional_correctness --help
 ```
 However, we recommend that you use the default values for the rest.
 
-## Known Issues
+## Issues
+Problem descriptions in `descriptions/VerilogDescriptions_Machine.jsonl` are machine 
+generated and we can not guarantee the absense of ambiguity and errors. We do not plan
+to maintain description correctness.
+
+Functional correctness are evaluated through comparing simulation outputs using 
+[ICARUS Verilog](https://github.com/steveicarus/iverilog). The evaluation of Verilog syntax is limited by the simulator, which might not include all features of Verilog HDL 
+IEEE-1364 standard.
 
-While evaluation uses very little memory, you might see the following error
-message when the system is running out of RAM. Since this may cause some
-correct programs to fail, we recommend that you free some memory and try again.
-```
-malloc: can't allocate region
-```
 
 ## Citation
 
 Please cite using the following bibtex entry:
 
 ```
-@article{chen2021codex,
-  title={Evaluating Large Language Models Trained on Code},
-  author={Mark Chen and Jerry Tworek and Heewoo Jun and Qiming Yuan and Henrique Ponde de Oliveira Pinto and Jared Kaplan and Harri Edwards and Yuri Burda and Nicholas Joseph and Greg Brockman and Alex Ray and Raul Puri and Gretchen Krueger and Michael Petrov and Heidy Khlaaf and Girish Sastry and Pamela Mishkin and Brooke Chan and Scott Gray and Nick Ryder and Mikhail Pavlov and Alethea Power and Lukasz Kaiser and Mohammad Bavarian and Clemens Winter and Philippe Tillet and Felipe Petroski Such and Dave Cummings and Matthias Plappert and Fotios Chantzis and Elizabeth Barnes and Ariel Herbert-Voss and William Hebgen Guss and Alex Nichol and Alex Paino and Nikolas Tezak and Jie Tang and Igor Babuschkin and Suchir Balaji and Shantanu Jain and William Saunders and Christopher Hesse and Andrew N. Carr and Jan Leike and Josh Achiam and Vedant Misra and Evan Morikawa and Alec Radford and Matthew Knight and Miles Brundage and Mira Murati and Katie Mayer and Peter Welinder and Bob McGrew and Dario Amodei and Sam McCandlish and Ilya Sutskever and Wojciech Zaremba},
-  year={2021},
-  eprint={2107.03374},
-  archivePrefix={arXiv},
-  primaryClass={cs.LG}
+@inproceedings{liu2023verilogeval,
+  title={{VerilogEval:} Evaluating Large Language Models for Verilog Code Generation},
+  author={Liu, Mingjie and Pinckney, Nathaniel and Khailany, Brucek and Ren, Haoxing},
+  booktitle={2023 IEEE/ACM International Conference on Computer-Aided Design (ICCAD)}, 
+  year={2023}
 }
 ```