This is default featured slide 1 title

Go to Blogger edit html and find these sentences.Now replace these sentences with your own descriptions.

This is default featured slide 2 title

Go to Blogger edit html and find these sentences.Now replace these sentences with your own descriptions.

This is default featured slide 3 title

Go to Blogger edit html and find these sentences.Now replace these sentences with your own descriptions.

This is default featured slide 4 title

Go to Blogger edit html and find these sentences.Now replace these sentences with your own descriptions.

This is default featured slide 5 title

Go to Blogger edit html and find these sentences.Now replace these sentences with your own descriptions.

Friday, August 25, 2023

New best story on Hacker News: Beating GPT-4 on HumanEval with a fine-tuned CodeLlama-34B

Beating GPT-4 on HumanEval with a fine-tuned CodeLlama-34B
418 by rushingcreek | 144 comments on Hacker News.
Hi HN, We have fine-tuned CodeLlama-34B and CodeLlama-34B-Python on an internal Phind dataset that achieved 67.6% and 69.5% pass@1 on HumanEval, respectively. GPT-4 achieved 67%. To ensure result validity, we applied OpenAI's decontamination methodology to our dataset. The CodeLlama models released yesterday demonstrate impressive performance on HumanEval. - CodeLlama-34B achieved 48.8% pass@1 on HumanEval - CodeLlama-34B-Python achieved 53.7% pass@1 on HumanEval We have fine-tuned both models on a proprietary dataset of ~80k high-quality programming problems and solutions. Instead of code completion examples, this dataset features instruction-answer pairs, setting it apart structurally from HumanEval. We trained the Phind models over two epochs, for a total of ~160k examples. LoRA was not used — both models underwent a native fine-tuning. We employed DeepSpeed ZeRO 3 and Flash Attention 2 to train these models in three hours using 32 A100-80GB GPUs, with a sequence length of 4096 tokens. Furthermore, we applied OpenAI's decontamination methodology to our dataset to ensure valid results, and found no contaminated examples. The methodology is: - For each evaluation example, we randomly sampled three substrings of 50 characters or used the entire example if it was fewer than 50 characters. - A match was identified if any sampled substring was a substring of the processed training example. For further insights on the decontamination methodology, please refer to Appendix C of OpenAI's technical report. Presented below are the pass@1 scores we achieved with our fine-tuned models: - Phind-CodeLlama-34B-v1 achieved 67.6% pass@1 on HumanEval - Phind-CodeLlama-34B-Python-v1 achieved 69.5% pass@1 on HumanEval Note on GPT-4 According to the official technical report in March, OpenAI reported a pass@1 score of 67% for GPT-4's performance on HumanEval. Since then, there have been claims reporting higher scores. However, it's essential to note that there hasn't been any concrete evidence pointing towards an enhancement in the model's coding abilities since then. It's also crucial to highlight that these elevated figures lack the rigorous contamination analysis that the official statistic underwent, making them less of a reliable comparison. As a result, we consider 67% as the pass@1 score for GPT-4. Download We are releasing both models on Huggingface for verifiability and to bolster the open-source community. We welcome independent verification of results. Phind-CodeLlama-34B-v1: https://ift.tt/u9eMyab Phind-CodeLlama-34B-Python-v1: https://ift.tt/DYXAbSZ We'd love to hear your thoughts! Best, The Phind Team

New best story on Hacker News: Web scraping for me, but not for thee

New best story on Hacker News: Factorio: Space Age

Factorio: Space Age
484 by haunter | 117 comments on Hacker News.


New best story on Hacker News: The complete sequence of a human Y chromosome

New best story on Hacker News: Hacker News Guidelines

Hacker News Guidelines
391 by tonmoy | 361 comments on Hacker News.


Wednesday, August 23, 2023

New best story on Hacker News: Common mistakes in salary negotiation

Common mistakes in salary negotiation
412 by eamonnm | 326 comments on Hacker News.


New best story on Hacker News: AI real-time human full-body photo generator

New best story on Hacker News: Ask HN: Where to find open-source house plans?

Ask HN: Where to find open-source house plans?
387 by tsingy | 192 comments on Hacker News.
Wanting to build a house, and looking for a DB of open source plans if such thing even exist.

New best story on Hacker News: Don't fire your illustrator

Don't fire your illustrator
374 by todsacerdoti | 316 comments on Hacker News.


New best story on Hacker News: I only lost 10 minutes of data, thanks to ZFS

I only lost 10 minutes of data, thanks to ZFS
379 by chromakode | 234 comments on Hacker News.


New best story on Hacker News: I walked across Luxembourg

I walked across Luxembourg
376 by shoobs | 216 comments on Hacker News.


Tuesday, August 22, 2023

Monday, August 21, 2023

New best story on Hacker News: uBlock Origin Lite now available on Firefox

New best story on Hacker News: John Warnock has died

John Warnock has died
557 by skilled | 124 comments on Hacker News.