takminの書きっぱなし備忘録 @はてなブログ

主にコンピュータビジョンなど技術について、たまに自分自身のことや思いついたことなど

2026/07/18第67回CV勉強会@関東「CVPR2026読み会」(前編)資料まとめ

7月18日は、第67回コンピュータビジョン勉強会@関東「CVPR2026読み会」(前編)を、株式会社DeNA/IRIAM様に会場をお借りして行いました。

以下、自分で見返すために資料やリンク等をまとめておきます。

登録サイト

kantocv.connpass.com

posfie

posfie.com

Youtube

今回は途中で映像が流れなかったり、音声が流れなかったりしてしまい、申し訳ございません。

www.youtube.com

コンピュータビジョン勉強会@関東

sites.google.com

資料まとめ

発表者 論文タイトル 発表資料
alfredplpl Back to Basics: Let Denoising Generative Models Denoise https://www.docswell.com/s/alfredplpl/Z8N989-2026-07-18-082234
shunk031 PSDesigner: Automated Graphic Design with a Human-Like Creative Workflow https://speakerdeck.com/shunk031/kantocv-67th-cvpr-2026
lychee1223_Lab SciGA: A Comprehensive Dataset for Designing Graphical Abstracts in Academic Papers
Paper2Figure: A Multi-Agent Collaborative System for Figure Generation Towards Academic Research Paper
https://speakerdeck.com/lychee1223/kantocv-67th-cvpr-2026
kzykmyzw SoccerMaster: A Vision Foundation Model for Soccer Understanding https://speakerdeck.com/kzykmyzw/soccermaster-a-vision-foundation-model-for-soccer-understanding
Ito What Is the Optimal Ranking Score Between Precision and Recall? We Can Always Find It and It Is Rarely F1 https://speakerdeck.com/keiichiito1978/20260618-cvprdu-mihui
welldone Sensor2Sensor: Cross-Embodiment Sensor Conversion for Autonomous Driving https://speakerdeck.com/welldone/cvmian-qiang-hui-sensor2sensor
nasnetou Can You Learn to See Without Images? Procedural Warm-Up for Vision Transformers https://docs.google.com/presentation/d/13WFHhwXBmyhoa1bDJyeU8CvrrTLb-yZAacYtwr3eiw4/
Kenji AT-VLA: Adaptive Tactile Injection for Enhanced Feedback Reaction in Vision-Language-Action Models https://speakerdeck.com/tsukamotokenji/at-vla-b09cef8c-8208-4d1c-82e2-6cf35ec1ac84
TatsuyaSuzuki RoboWheel: A Data Engine from Real-World Human Demonstrations for Cross-Embodiment Robotic Learning https://speakerdeck.com/x_ttyszk/di-67hui-konpiyutabiziyonmian-qiang-hui-lun-wen-shao-jie-robowheel-a-data-engine-from-real-world-human-demonstrations-for-cross-embodiment-robotic-learning

後編は8/29(土)に(株)エクサウィザーズ様の会場をお借りして開催予定です。

kantocv.connpass.com

2026/02/08第66回CV勉強会@関東「世界モデル論文読み会」資料まとめ

2月8日は、第66回コンピュータビジョン勉強会@関東「世界モデル論文読み会」を、チューリング株式会社様に会場をお借りして行いました。

以下、自分で見返すために資料やリンク等をまとめておきます。

登録サイト

kantocv.connpass.com

posfie

posfie.com

Youtube

今回はネットワークが不安定だったため、動画が途切れ途切れになってしまい、申し訳ございません。

www.youtube.com

www.youtube.com

www.youtube.com

コンピュータビジョン勉強会@関東

sites.google.com

資料まとめ

発表者 論文タイトル 発表資料
tomoaki_teshima World Models (2018) / Dreamer (2019) Dropbox - kantocv_world_models.pdf - Simplify your life
Keiichi-Ito WorldPrediction: A Benchmark for High-level World Modeling and Long-horizon Procedural Planning 20260208_第66回 コンピュータビジョン勉強会 - Speaker Deck
Hidehisa Arai 世界モデルにおける分布外データ対応の方法論 世界モデルにおける分布外データ対応の方法論 - Speaker Deck
Godel 自律移動ロボットはWorld Modelsの夢は見るか?
大政孝充 GWM: Towards Scalable Gaussian World Models for Robotic Manipulation 第66回コンピュータビジョン勉強会@関東における株式会社ウェブファーマー 大政の発表 GWMモデル | PPTX
caprest 世界モデルで物を掴めるようになるのか? 世界モデルで物を掴めるようになるのか?.pdf - Google ドライブ
Kento Sasaki Epona: Autoregressive Diffusion World Model for Autonomous Driving (ICCV 2025) 第66回コンピュータビジョン勉強会@関東 Epona: Autoregressive Diffusion World Model for Autonomous Driving - Speaker Deck
takmin Cosmos World Foundation Model Platform for Physical AI Cosmos World Foundation Model Platform for Physical AI - Speaker Deck
Shin-kyoto Alpamayo-R1: Bridging Reasoning and Action Prediction for Generalizable Autonomous Driving in the Long Tail [CV勉強会@関東 世界モデル論文読み会] VLA自動運転model Alpamayo-R1 - Speaker Deck
abemii Orbis: Overcoming Challenges of Long-Horizon Prediction in Driving World Models [CV勉強会@関東 World Model 読み会] Orbis: Overcoming Challenges of Long-Horizon Prediction in Driving World Models (Mousakhan+, NeurIPS 2025) - Speaker Deck

takmin発表資料

speakerdeck.com

雪であまり人が集まらないかと心配しましたが、とても盛況で安心しました。

2025/11/16第65回CV勉強会@関東「ICCV2025読み会」資料まとめ

11月16日は、第65回コンピュータビジョン勉強会@関東「ICCV2025読み会」を、渋谷スクランブルスクエアの株式会社MIXI様に会場をお借りして行いました。

以下、自分で見返すために資料やリンク等をまとめておきます。

登録サイト

kantocv.connpass.com

Togetterあらためposfie

posfie.com

YouTube

www.youtube.com

コンピュータビジョン勉強会@関東

sites.google.com

資料まとめ

発表者 発表内容 資料
abemii World4Drive: End-to-End Autonomous Driving via Intention-aware Physical Latent World Model https://speakerdeck.com/abemii/cvmian-qiang-hui-at-guan-dong-cvpr2025-du-mihui-world4drive-end-to-end-autonomous-driving-via-intention-aware-physical-latent-world-model-zheng-plus-iccv-2025
caprest Fine-Grained Evaluation of Large Vision-Language Models in Autonomous Driving https://docs.google.com/presentation/d/10JJiQdOwoKiw--6g7B5Mc2N2Y4BMpNKDCkQ5oSx5NZA
Shin End-to-End Driving with Online Trajectory Evaluation via BEV World Model https://speakerdeck.com/shinkyoto/cvmian-qiang-hui-at-guan-dong-iccv2025-wote-end-to-end-driving-with-online-trajectory-evaluation-via-bev-world-model
Keiichi-Ito AnimeGamer: Infinite Anime Life Simulation with Next Game State Prediction https://speakerdeck.com/keiichiito1978/20250916-di-65hui-konpiyutabiziyonmian-qiang-hui
alfredplpl ICCV 2025の動画生成を眺めてみた
StreamDiffusion: A Pipeline-level Solution forReal-Time Interactive Generation
VACE: All-in-OneVideo Creation and Editing
FiVE: A Fine-grainedVideo EditingBenchmark for EvaluatingEmerging Diffusion and Rectified Flow Models
https://www.docswell.com/s/alfredplpl/5VM392-2025-11-08-194529
ShuN Makino Is CLIP ideal? No. Can we fix it? Yes! https://speakerdeck.com/shun6211/lun-wen-shao-jie-is-clip-ideal-no-can-we-fix-it-yes-di-65hui-konpiyutabiziyonmian-qiang-hui-at-guan-dong
Kenji CoA-VLA: Improving Vision-Language-Action Models via Visual-Textual Chain-of-Affordance https://speakerdeck.com/tsukamotokenji/di-65hui-konpiyutabiziyonmian-qiang-hui
peisuke Moto: Latent Motion Token as the Bridging Language for Learning Robot Manipulation from Videos https://speakerdeck.com/peisuke/moto-latent-motion-token-as-the-bridging-language-for-learning-robot-manipulation-from-videos

今回、世界モデルに関する発表が多かったです。コンピュータビジョン業界の大きなトレンドになっているのかな?

2025/08/24第64回CV勉強会@関東「CVPR2025読み会」(後編)資料まとめ

8月24日は、7月13日の前編に引き続き、第64回コンピュータビジョン勉強会@関東「CVPR2025読み会」後編を、渋谷スクランブルスクエアの株式会社ディー・エヌ・エー様/株式会社IRIAM様に会場をお借りして行いました。

以下、自分で見返すために資料やリンク等をまとめておきます。

登録サイト

kantocv.connpass.com

Togetterあらためposfie

posfie.com

YouTube

www.youtube.com

コンピュータビジョン勉強会@関東

sites.google.com

資料まとめ

発表者 発表内容 資料
takmin R-SCoRe: Revisiting Scene Coordinate Regression for Robust Large-Scale Visual Localization https://speakerdeck.com/takmin/r-score-revisiting-scene-coordinate-regression-for-robust-large-scale-visual-localization
Takeo Shibata MotionPro: A Precise Motion Controller for Image-to-Video Generation
abemii MegaSaM: Accurate, Fast, and Robust Structure and Motion from Casual Dynamic Videos https://speakerdeck.com/abemii/cvmian-qiang-hui-at-guan-dong-cvpr2025-du-mihui-megasam-accurate-fast-and-robust-structure-and-motion-from-casual-dynamic-videos-li-plus-cvpr2025
s_aiueo32 Linguistics-aware Masked Image Modeling for Self-supervised Scene Text Recognition https://speakerdeck.com/s_aiueo32/cvpr2025lun-wen-du-mihui-linguistics-aware-masked-image-modeling-for-self-supervised-scene-text-recognition
kzykmyzw Gaze-LLE: Gaze Target Estimation via Large-Scale Learned Encoders https://speakerdeck.com/kzykmyzw/gaze-lle-gaze-target-estimation-via-large-scale-learned-encoders
Kenji RoboSpatial: Teaching Spatial Understanding to 2D and 3D Vision-Language Models for Robotics https://speakerdeck.com/tsukamotokenji/di-64hui-konpiyutabiziyonmian-qiang-hui-at-guan-dong-hou-bian
caprest SimLingo: Vision-Only Closed-Loop Autonomous Driving with Language-Action Alignment https://docs.google.com/presentation/d/1YqppSVFJNaqXKHuhx3SZvob1ZCulix2X-o8PWM7Zfao/
frkake Removing Reflections from RAW Photos https://speakerdeck.com/frkake/removing-reflections-from-raw-photos
YutaKikuchi DroneSplat: 3D Gaussian Splatting for Robust 3D Reconstruction from In-the-Wild Drone Imagery https://speakerdeck.com/yutakikuchi_sd/cvmian-qiang-hui-at-guan-dong-dronesplat-3d-gaussian-splatting-for-robust-3d-reconstruction-from-in-the-wild-drone-imagery
Keiichi-Ito Towards Zero‑Shot Anomaly Detection and Reasoning with Multimodal Large Language Models https://speakerdeck.com/keiichiito1978/cvprmian-qiang-hui-hou-ban

今回は私も発表したので、こちらに発表資料を埋め込んでおきます。

speakerdeck.com