pl.aiwright - GPT-4 dialogue for Disco Elysium: The Final Cut

nsa@kbin.social · 10 months ago

pl.aiwright - GPT-4 dialogue for Disco Elysium: The Final Cut

nsa@kbin.social · 1 year ago

It seems like for creative text generation tasks, metrics have been shown to be deficient; this even holds for the new model-based metrics. That leaves human evaluation (both intrinsic and extrinsic) as the gold standard for those types of tasks. I wonder if the results from this paper (and other future papers that look automatic CV metrics) will lead reviewers to demand more human evaluation in CV tasks like they do for certain NLP tasks.

nsa@kbin.social · 1 year ago

Exposing flaws of generative model evaluation metrics and their unfair treatment of diffusion models

nsa@kbin.social · 1 year ago

hmmm… not sure which model you’re referring to. do you have a paper link?

nsa@kbin.social · 1 year ago

do you have a link?

nsa@kbin.social · 1 year ago

@Koffindodjer indeed you are!

nsa@kbin.social · 1 year ago

Extending Context Window of Large Language Models via Positional Interpolation

nsa@kbin.social · 1 year ago

Inverse Scaling: When Bigger Isn't Better

nsa@kbin.social · 1 year ago

Craft an Iron Sword: Dynamically Generating Interactive Game Characters by Prompting Large Language Models Tuned on Code

nsa@kbin.social · 1 year ago

r/MachineLearning finally received a warning from u/ModCodeOfConduct

nsa@kbin.social · 1 year ago

If the effect is strong enough, then it could have a very negative effect on LLM training in the near future, considering more and more of the internet contains ChatGPT & GPT-4 content in it and automatic detectors are currently quite poor.

nsa